Prepare for Linux support interviews by practicing the commands, concepts, and troubleshooting scenarios employers expect you to understand, not just recite.
Real questions, real answers, the kind you could actually say out loud in an interview. Click any question to expand it.
chmod 755 mean, and when would you use it versus 644?Whether you actually understand the permission model rather than memorizing numbers.
Permissions come in three groups, owner, group, and everyone else, and each gets read (4), write (2), execute (1). So 755 is rwx for the owner and r-x for group and others. I would use 755 on a directory or a script that needs to be executable, and 644 (rw-r--r--) on a regular file like a config or a document, because normal files should not be executable.
The reason the execute bit matters on directories is that without it you cannot cd into the directory even if you can read its name. So a common real-world bug is a directory set to 644 that people can see but cannot enter. When I set permissions I think about who actually needs access and give the least that works, rather than reaching for 777, which is almost always a sign someone gave up.
Reaching for chmod 777 to “make it work.” It removes all protection and is a red flag to an experienced interviewer.
“On a directory, what does the execute bit actually let you do?”
kill -15 and kill -9?Whether you understand signals and default to the graceful option.
kill -15 sends SIGTERM, which is a polite request asking the process to shut down and clean up after itself. kill -9 sends SIGKILL, which the process cannot catch or ignore, the kernel just ends it. I try SIGTERM first and only use SIGKILL if the process is stuck and not responding.
The reason order matters is that SIGKILL skips cleanup, so a database or a process holding a lock can be left in a bad state, corrupt files or stale lock files. My habit is: find the process, send SIGTERM, wait a few seconds, confirm it is gone, and only escalate to -9 if it ignored me. Reaching straight for -9 is a small tell that someone has not been burned by it yet.
Using kill -9 first by reflex.
“You send SIGTERM and nothing happens. What do you check before escalating?”
Troubleshooting method and whether you read errors instead of guessing.
I start with systemctl status <service>, which shows whether it failed and usually a hint, then journalctl -u <service> for the actual error. A lot of services also have a config test command, like nginx -t, that points straight at the bad line. I fix the specific cause, validate the config again, then restart and confirm it is running.
Because a config change is the stated cause, I would also check whether there is a backup of the old config to compare against or roll back to, that is exactly why I back up a config before editing it. And I would not just restart blindly, I would run the config test first so I am not fixing one typo and introducing another. Then I confirm from the outside too, for a web server I would curl it, not just trust the status line.
Restarting repeatedly without reading the log, or editing config with no backup to roll back to.
“The log says the port is already in use. Now what?”
Structured investigation and awareness of the non-obvious causes.
I start with df -h to confirm which filesystem is full, then I narrow down with du, usually du -h --max-depth=1 /var and then drill into the biggest directory, and again, until I find the specific files. Very often it is a runaway log in /var/log.
A couple of things I check that catch people out. On a busy production box I never run du unscoped from /, without the -x flag it will cross into /proc, /sys, and any mounted network shares, which throws permission noise and can hang on a slow NFS mount. Adding -x keeps it on the local filesystem. Separately, if df -h shows space free but writes still fail, I check df -i, because the filesystem can run out of inodes from millions of tiny files. And if I find a huge log a process still has open, deleting it does not free the space until the process restarts, so I truncate it in place and fix the cause instead.
Only checking df and giving up when it shows free space, missing inode exhaustion.
“You delete the big log file but df still shows the space used. Why?”
The complete Linux Interview Prep Pack adds dozens more questions, deeper troubleshooting scenarios, follow-up questions, a quick-review cheat sheet, and a 7-day study plan.
Practical network-service troubleshooting.
I use ss -tlnp, which lists TCP ports that are listening and the process that owns each one. To check one port I pipe it through grep. This comes up constantly when a service is up but nothing can connect to it.
The detail that matters is which address it is bound to. A service listening on 127.0.0.1 only accepts local connections, so if a remote client cannot reach it but the service is clearly running, that binding is often the reason, and the fix is config, not firewall. ss replaced netstat, though netstat -tlnp still works on older boxes.
Assuming “the service is running” means “the service is reachable,” without checking the bind address.
“The service is listening on 127.0.0.1 but another server cannot reach it. What do you change?”
Log-analysis fluency, the daily reality of Linux support.
grep finds lines that match a pattern, so for logs I use it constantly, like grep "error" app.log or grep -i "failed" /var/log/auth.log. I reach for awk when I need to pull out a specific field rather than a whole line, for example the IP address or the timestamp column.
A pattern I use a lot is counting by field. To find the busiest source IPs in an access log I would do awk '{print $1}' access.log | sort | uniq -c | sort -rn | head. Read right to left, that is: take the first field, group identical ones, count them, and show the biggest. That one line turns a huge log into an answer, which is most of what log analysis actually is.
Running uniq without sorting first, so duplicates that are not adjacent do not get collapsed.
“How would you get just the failed SSH logins and count them by source IP?”
Secure remote administration, a core Linux support skill.
I generate a key pair with ssh-keygen, keep the private key on my machine, and put the public key in the server's ~/.ssh/authorized_keys, usually with ssh-copy-id. After that I can log in without typing a password. It is more secure because nothing guessable travels over the network.
On a server I control, once key login works I would disable password authentication entirely in sshd_config, which shuts down the constant password brute-force attempts you see in the auth log. One gotcha worth mentioning: SSH refuses to use keys if the permissions on ~/.ssh or authorized_keys are too open, so if key login mysteriously fails, ssh -v and the file permissions are the first things I check.
Committing or sharing the private key, or disabling password auth before confirming key login works and locking yourself out.
“You set up your key but still get 'Permission denied (publickey)'. What do you check?”
Foundational understanding, and whether you can explain a concept crisply.
The kernel is the core that talks to the hardware and manages memory, processes, and devices. A distribution is the kernel plus everything that makes it usable, a package manager, system utilities, a shell, and sensible defaults. Ubuntu, Debian, and Red Hat are distributions built on the same kernel lineage.
Why it matters day to day is that distributions differ in ways that bite you, different package managers (apt versus dnf), different default log locations, different service names. So when I am handed an unfamiliar server, one of the first things I check is which distribution and version it is, because it changes which commands and paths I reach for.
Treating all Linux as Ubuntu and assuming Ubuntu paths and commands work everywhere.
“You are on an unfamiliar server. How do you find out what it is running?”
These questions come from the free SecureByDefault curriculum. Go learn the underlying skills, or explore the scenarios these answers are built on.