Prepare for junior sysadmin interviews by practicing structured incident response and everyday administration, the scenario-heavy questions current hiring guides weight most.
Real questions, real answers, the kind you could actually say out loud in an interview. Click any question to expand it.
Calm, structured incident response, the most heavily weighted sysadmin question.
First I confirm the scope: is the whole server down or just one service, is it one server or several, and can I reach it at all. If I can get in, I check the basics fast, is it out of disk or memory, is the service running, and what do the logs say around the time it failed. I fix the immediate problem to restore service, then look at root cause once it is stable.
The instinct that matters is restore first, understand fully second, but capture evidence before I destroy it. So before I reboot anything, I grab what state I can, because a reboot often clears the very thing that would have told me why. I also communicate: post that I am on it and give an ETA, because during an outage silence is its own problem. Recent changes are my first hypothesis.
Rebooting immediately to make it go away, losing the evidence and the root cause with it.
“It comes back after a reboot but you never found the cause. Are you done?”
Everyday account administration and attention to a genuinely dangerous flag detail.
I use usermod -aG <group> <user>. The important part is the -a, which means append. If you leave it off and just do usermod -G, you replace all of the user's secondary groups with the one you named, which can quietly remove them from groups they needed.
That missing -a is one of those mistakes that does not error out, it just silently does the wrong thing, and you find out later when someone has lost access. I also remember that group changes do not take effect until the user starts a new session, so I verify with id <user> afterward rather than assuming it worked.
Omitting -a and overwriting the user's group memberships; also forgetting a fresh login is needed.
“You added the user to the group but they still get permission denied. Why?”
Whether you treat backups as a tested process, not a checkbox.
I make sure important data is backed up on a schedule, kept somewhere separate from the system it came from, and retained long enough to be useful. The part people get wrong is never testing the restore, so I actually restore from a backup periodically to confirm it works.
The line I live by is that a backup you have never restored is not a backup, it is a hope. Plenty of real outages became disasters because backups had been silently failing for weeks and nobody checked. I monitor that backups are actually completing, and I do a test restore on purpose, because the worst time to discover a broken backup is during the incident where you need it.
Assuming backups work because the job exists, without ever testing a restore or monitoring for failures.
“How would you know today whether your backups have been working for the last month?”
Balancing security with stability, a core sysadmin judgment.
I keep systems patched because unpatched software is one of the easiest ways in, but I do not just push updates blindly to everything at once. I test on a non-critical group first, then roll out more broadly, and I schedule it to minimize disruption.
The tension is real: patch too slowly and you are exposed, patch too fast and you can break production. So I stage it, test group, then production in waves, with a way to roll back. For urgent security issues I move faster but still watch the first wave closely. Automating and centralizing patching matters at scale, doing it by hand does not survive a large fleet.
Either patching everything at once with no testing, or letting patching slip because it is disruptive.
“A critical vulnerability drops and there is no time for full testing. How do you handle the risk?”
The complete Junior Sysadmin Interview Prep Pack adds dozens more questions, deeper troubleshooting scenarios, follow-up questions, a quick-review cheat sheet, and a 7-day study plan.
Permissions reasoning applied to a real, common ticket.
Since teammates can and this one user cannot, it is specific to them, so I check group membership first. I look at the directory's owner, group, and permissions, then check whether this user is actually in the group that has access.
Almost always the user is missing from the group the directory grants access to, so they fall through to “other,” which has no access. The fix is adding them with usermod -aG and having them re-log so it takes effect. I also check the directory has execute permission for the group, since without it you cannot enter even if you can read.
Changing the file permissions broadly instead of fixing the one user's group membership.
“Everyone including this user is in the right group, but they still cannot enter the directory. What now?”
Basic operational fluency with the right tools.
For a quick overall view, top or htop. For memory specifically, free -h gives a clean readable output. For disk, df -h shows filesystem usage and du -sh shows what is using space. uptime gives load at a glance.
On a virtual machine I also watch for CPU steal in top, the 'st' value, because high steal means the host is oversubscribed and the fix is not on my server at all. I also think about monitoring proactively rather than only checking after a page, the goal is to catch the disk filling up before it takes the service down.
Only checking reactively after an incident, with no monitoring to catch problems early.
“Load is high but CPU usage looks low. What could explain that?”
Debugging scheduled tasks, a classic sysadmin gotcha.
I check that the cron entry is actually there and correct with crontab -l, then check the logs to see if cron tried to run it. A lot of the time the job runs but fails, so I look at whether the script works when I run it by hand.
The most common real cause is environment: cron runs with a minimal environment and a different PATH than my interactive shell, so a script that works when I run it can fail under cron because a command is not found or a relative path is wrong. So I use full paths in cron jobs and redirect output to a log file so I can actually see the errors.
Assuming the schedule is wrong when the real issue is the cron environment or a relative path.
“The script runs perfectly by hand but fails under cron. What is the usual reason?”
A security-first mindset applied to setup, without overcomplicating it.
I update the system first so I start patched, create a normal user instead of working as root, set up SSH key authentication and lock down SSH, enable a firewall allowing only what is needed, and make sure I can see the logs.
I sequence it so I never lock myself out: get key login working and confirm it before I disable password auth. Then firewall to only the required ports, disable direct root login, and set up basic monitoring and backups. I am not building a hardened bastion on day one, I am closing the doors that automated attacks walk through, and those attacks start within hours of a server going live.
Disabling password login before confirming key login works, and locking yourself out of a remote server.
“Why confirm key login works before you disable password authentication?”
These questions come from the free SecureByDefault curriculum. Go learn the underlying skills, or explore the scenarios these answers are built on.