PRACTICE TRACK / 118 QUESTIONS

Linux
Think it through.

Processes, signals, networking, and troubleshooting.

Choose a question, explain your approach, then reveal the supplied answer. Difficulty labels come from the existing question library.

118 questions

Answers stay closed until you choose to reveal them.

QUESTION 01LinuxEasy

What is the difference between a hard link and a symbolic link?

#
Reveal answer guidance

A hard link is another directory entry pointing to the same inode — the file exists as long as any hard link remains, and all links are equal. A symlink is a small file that stores a path to another file; it can cross filesystems and link directories, but breaks if the target is moved or deleted. Hard links can't cross filesystems or point to directories.

QUESTION 02LinuxEasy

How do you find which process is using a port?

#
Reveal answer guidance

Use ss -ltnp 'sport = :8080' (or lsof -i :8080, or older netstat -ltnp) to list the listening socket and its PID/program. Then inspect or kill it with ps -p <pid> / kill. ss is the modern, faster replacement for netstat.

QUESTION 03LinuxMedium

A server is slow. Walk through how you triage it.

#
Reveal answer guidance

Start broad with top/htop and uptime (load average vs core count). Then split by resource: CPU (top, pidstat), memory (free -h, check swap and OOM in dmesg), disk I/O (iostat -x, iotop), and network (ss, iftop). vmstat 1 gives a quick overall view. The goal is to localize the bottleneck to CPU, memory, I/O, or network before drilling into the offending process or query.

QUESTION 04LinuxMedium

What do the permission bits 755 and the SUID bit mean?

#
Reveal answer guidance

755 = owner rwx (7), group r-x (5), others r-x (5) — common for executables and directories. SUID (e.g. 4755, the leading 4) makes an executable run with the file owner's privileges rather than the caller's — that's how passwd edits /etc/shadow as root. SUID on binaries is a security-sensitive setting and a common privilege-escalation vector.

QUESTION 05LinuxHard

What is the difference between SIGTERM, SIGKILL, and SIGHUP?

#
Reveal answer guidance

SIGTERM (15) politely asks a process to terminate — it can catch it, clean up, and exit; this is what you should use first and what kill sends by default. SIGKILL (9) cannot be caught or ignored; the kernel kills the process immediately with no cleanup — last resort. SIGHUP originally meant terminal hangup but by convention tells daemons to reload their config without restarting. Graceful shutdown handlers in apps/containers hook SIGTERM.

QUESTION 06LinuxHard

A process inside a Docker container (which uses cgroups v2) reports memory usage of 4GB via free -h, but the host shows the container's cgroup memory limit is 2GB. What is happening, and how does cgroup v2 memory accounting differ from the kernel's view of the process?

#
Reveal answer guidance

free -h reads from /proc/meminfo, which shows the host's global memory statistics, not the container's cgroup limits. In cgroup v2, the container's memory limit is enforced at the cgroup level, not via proc. The correct way to see container memory usage is cat /sys/fs/cgroup/memory.current and cat /sys/fs/cgroup/memory.max. The 4GB shown by free is the host's total memory. cgroup v2 memory accounting includes: memory.current (total page cache + RSS + swap), memory.swap.current (swap usage), and memory.stat for detailed breakdown (anon, file, kernel, slab). Unlike cgroup v1, v2 unifies memory and memsw accounting — the memory.max limit applies to memory+swap combined. The kernel enforces OOM via memory.events — when max is exceeded, oom_kill increments. If free shows 4GB, the process is seeing host memory. For accurate container-level memory, always read the cgroup files. Docker's docker stats reads these cgroup values correctly.

QUESTION 07LinuxHard

A systemd service fails to start with "Unit exited due to start timeout". The ExecStart command runs a Python script that sometimes takes 5 minutes. How do you reconfigure systemd without changing the script, and what's the default timeout?

#
Reveal answer guidance

The default TimeoutStartSec is 90 seconds (systemd v240+). When the script exceeds this, systemd sends SIGTERM (configurable via KillSignal), then SIGKILL after TimeoutStopSec. To fix: override the timeout in a drop-in via systemctl edit myservice.service and add [Service] TimeoutStartSec=300 (or infinity). However, if the script hangs indefinitely, this masks the problem. Better: use Type=notify with sd_notify("READY=1") in the script, so systemd knows the service is alive without waiting for the full timeout. If the script is not modifiable, use Type=simple with ExecStartPost that polls for readiness — or use Restart=on-failure with RestartSec=30 to handle transient delays. Also consider TimeoutStartSec behavior differs between Type=simple (systemd waits for the execve only) and Type=oneshot or Type=notify (waits for process exit or notification). For long initialization, Type=idle delays start until boot completes.

QUESTION 08LinuxMedium

A developer runs strace -p <PID> on a stuck process and sees futex(0x7f..., FUTEX_WAIT_PRIVATE, 2, NULL) repeating. What does this mean, and how do you identify which mutex is contended?

#
Reveal answer guidance

futex(FUTEX_WAIT_PRIVATE) means the thread is blocked waiting on a userspace mutex. The futex syscall is the kernel-level wait operation — when FUTEX_WAIT_PRIVATE returns EWOULDBLOCK, the thread re-checks the userspace lock word; if still contended, it sleeps. The address 0x7f... is the userspace address of the futex word (part of the pthread_mutex_t or pthread_cond_t). To identify the contended mutex: (1) Check /proc/<PID>/maps to find which shared library or heap region maps that address. (2) Run gdb -p <PID>, then info threads, thread apply all bt to see which thread holds the lock. (3) Use futex traces from all threads — the lock owner is not in FUTEX_WAIT but in FUTEX_WAKE or running code. (4) Enable lockdep with echo 1 > /sys/kernel/debug/tracing/events/lock/lock_acquire/enable. (5) Use perf lock record -p <PID> then perf lock report to show contended locks by symbol. For Go programs, use GOTRACEBACK=all or the runtime/trace output instead.

QUESTION 09LinuxHard

Explain how Linux namespaces interact with cgroups to provide container isolation. Specifically, how does a user namespace map root (UID 0) inside the container to an unprivileged UID on the host, and what are the security implications?

#
Reveal answer guidance

User namespaces map a range of UIDs inside the namespace to a different range on the host via /proc/<PID>/uid_map. For example, 0 100000 65536 maps container UID 0 to host UID 100000. This means the container's root has no special privileges on the host — kernel capabilities are scoped to the namespace. When a process in the container tries to mknod or setcap, the kernel checks the mapped host UID, not the namespace UID. Without user namespaces, containers share the host's UID space — if allow userns=0 isn't enabled, an escape from a root-in-container process becomes real root on the host. However, user namespaces add complexity: (1) Mounting filesystems requires --privileged or CAP_SYS_ADMIN in the initial namespace. (2) The subuid and subgid files (in /etc/) define the mapping ranges used by tools like Podman. (3) Not all kernel interfaces are namespace-aware — for example, /proc/sysrq-trigger is global. Docker by default blocks user namespace remapping for networking (since netdev ownership is global), so rootless containers in Docker require extra networking setup (slirp4netns).

QUESTION 10LinuxHard

You run a production database on ext4. Write-heavy workloads cause jbd2 process to consume high CPU and I/O wait spikes. Diagnose the root cause and explain if XFS would behave differently.

#
Reveal answer guidance

jbd2 is the ext4 journaling daemon. High CPU/IO means the journal is a bottleneck — every metadata write goes to the journal first (write-ahead logging). ext4 defaults to data=ordered mode: data blocks are written before metadata is journaled. Under high write concurrency, the journal commits become serialized, and jbd2 flushes the journal to disk at each commit (default 5-second commit interval, or sooner if the journal fills). XFS uses a different design: it allocates log space in a circular buffer, groups multiple metadata updates into a single transaction, and uses delayed logging — dirty metadata is accumulated in memory and committed asynchronously. XFS typically handles concurrent metadata writes better because it batches more aggressively and its log writes use asynchronous I/O when the storage supports it. However, both filesystems will show jbd-like activity under extreme write pressure. Mitigations for ext4: increase the journal size (128MB default), place journal on a separate SSD/NVMe device, use barrier=0 if you accept crash-risk, or switch to data=writeback (metadata-only journaling, risk of data corruption on crash).

QUESTION 11LinuxMedium

A server runs with swap enabled (swappiness = 60). The application starts showing high latency even though free -h shows 70% RAM free. vmstat shows si and so are non-zero. What's happening, and how do you fix it?

#
Reveal answer guidance

This is premature swapping. The kernel's page reclaim logic may swap out inactive anonymous pages even when plenty of file-backed page cache is available. Swappiness (default 60) controls the balance between reclaiming file-backed pages vs anonymous pages — at 60, the kernel is slightly biased to reclaim file-backed pages. If the working set includes many mapped pages (e.g., Java heap, mmap'd files), the kernel may swap them to avoid evicting cached files. The fix: sudo sysctl vm.swappiness=1 (or 0 on kernels < 5.8+ where 0 disables swap entirely). Also check vm.vfs_cache_pressure — lowering it retains dentry/inode caches. Additional diagnostics: sar -S to see swap rates, cat /proc/meminfo | grep -i swap to find what's swapped (use smem or swap from procps-ng). If the application is Java, verify the JVM heap fits in RAM and consider -XX:+UseTransparentHugePages to reduce TLB misses. For databases, disable swap entirely (swapoff -a) or set swappiness=1 and configure vm.overcommit_memory=2 to prevent memory-overcommit-based swapping.

QUESTION 12LinuxMedium

What is the practical difference between SELinux's targeted and strict policies, and how would you troubleshoot a process that gets "AVC denied" even though the file permissions (rwx) are correct?

#
Reveal answer guidance

targeted applies SELinux policy only to specific "targeted" daemons (httpd, sshd, named) while leaving other processes unconfined. strict applies policy to all processes, including user sessions — rarely used in production because it breaks almost everything unless the policy is custom-tailored. An "AVC denied" with correct rwx means the file's SELinux context (user:role:type:level) does not match the process's domain. For example, httpd runs in httpd_t domain, but the file is labeled user_home_t (default for files in /home/). Fix: chcon -t httpd_sys_content_t /path/to/file or semanage fcontext -a -t httpd_sys_content_t /path(/.*)? then restorecon -Rv /path/. To troubleshoot: run ausearch -m avc -ts recent | audit2why to decode the denial. audit2allow generates a local policy module to permit the access, but in production, you should instead set the correct context. For debugging temporarily: setenforce 0 (permissive mode) — but never leave it disabled in production.

QUESTION 13LinuxHard

A systemd timer triggers a service every 5 minutes, but sometimes the service takes 7 minutes. The timer starts a new instance anyway, causing overlapping runs. How do you prevent overlap — using flock in the script vs using systemd's unit options?

#
Reveal answer guidance

In the service unit: set [Service] Type=simple with no concurrency control — systemd will start a new instance whenever the timer fires. To prevent overlap, use [Unit] Conflicts=myservice.service combined with [Service] RemainAfterExit=yes — but this is convoluted. The idiomatic systemd solution: in the timer unit, set Unit=myservice.service without OnCalendar only — that doesn't prevent overlap. Instead, configure the service to be Type=oneshot and add [Service] StartLimitIntervalSec=0 with Restart=no, and in the timer, set OnCalendar=*:0/5 with AccuracySec=1s — but this still allows overlap. The correct approach: set [Unit] in the service to StopWhenUnneeded=no and use [Service] ExecStartPre=/usr/bin/flock -n /var/lock/myservice.lock. Better: add [Unit] ConditionPathExists=!/var/lock/myservice.lock and have the script create the lock. The cleanest systemd-native approach: use a [Unit] with BindsTo=myservice.service and After=myservice.service but that's tricky with timers. The practical answer: use flock -n /tmp/myservice.lock /path/to/script in ExecStart — the -n exits immediately if lock is held, and FailureAction= in the timer can handle errors.

QUESTION 14LinuxMedium

A process is using 100% CPU but you cannot tell from top which code path is hot. What tools beyond top do you use to identify the exact function in a compiled binary?

#
Reveal answer guidance

top shows aggregate CPU but not per-function breakdown. Use perf top -p <PID> to see the hottest functions in real-time, with symbol names. For JIT-compiled code (Node, JVM, .NET), use perf top -p <PID> -k 0 (to skip kernel symbols) with -g for call-graph. If symbols are missing (stripped binary), install debug symbols or use perf top -f for raw instruction-pointer sampling. Other tools: flamegraph via perf record -F 99 -g -p <PID> -- sleep 30 && perf script | stackcollapse-perf.pl | flamegraph.pl > out.svg. For Python: py-spy or cProfile. For Java: async-profiler. For short bursts, use perf stat -p <PID> -e cycles,instructions,cache-misses,branch-misses -- sleep 10 to see microarchitectural bottlenecks, then drill with perf record and perf annotate. If the binary is compiled with -fno-omit-frame-pointer, call graphs are accurate; without it, perf uses DWARF unwinding (--call-graph dwarf).

QUESTION 15LinuxHard

You are on a system without cron installed, and you cannot install packages. Only systemd is available. Write a persistence mechanism that runs a cleanup script every 6 hours without a timer unit file?

#
Reveal answer guidance

You can create a transient systemd timer using systemd-run. The command: systemd-run --user --on-calendar=daily --unit=cleanup /path/to/script.sh — but for 6-hour intervals: systemd-run --on-active=6h --unit=cleanup6h /path/to/script.sh. However, --on-calendar requires proper calendar syntax. For a repeating 6-hour schedule: systemd-run --on-calendar="*:0/360" --unit=cleanup /path/to/script.sh. If you cannot use --on-calendar, use systemd-run --unit=cleanup.service --timer-property=OnUnitActiveSec=6h /path/to/script.sh. This creates a transient timer and service unit in /run/systemd/transient/ that persists until next reboot. For reboot persistence without package installs: write a service unit to /etc/systemd/system/ directly (no systemctl daemon-reload needed if you use systemctl link). If even file write is restricted, use systemd-run --user --scope to create a scope, but it won't survive reboot. The --user instance requires lingering enabled (loginctl enable-linger). The --on-active creates a monotonic timer that runs once — for repeating, use OnUnitActiveSec in the timer properties.

QUESTION 16LinuxMedium

An NFS mount on your Linux server causes ls -la to hang for 30+ seconds when the NFS server is slow. How do you diagnose which mount is hanging and fix it without rebooting?

#
Reveal answer guidance

Diagnose: (1) Run mount | grep nfs to list all NFS mounts. (2) Run cat /proc/mounts | grep nfs and check the soft vs hard option — hard mounts block I/O indefinitely if the server is unreachable. (3) Use strace -p <PID_of_ls> to see if it's stuck on stat or getdents on an NFS path. (4) Run nfsstat -m for per-mount statistics. (5) Quick fix for a hanging mount: umount -f -l /path/to/mount (force lazy unmount) — the filesystem is detached immediately, processes get I/O errors and resume. (6) For a soft recovery: mount -o remount,soft,timeo=100,retrans=3,retry=0 /path — soft causes operations to timeout and return errors instead of hanging; timeo (tenths of seconds) sets the retransmission timeout. (7) Long-term: use soft,intr,timeo=600 for NFSv3 or soft,timeo=600 for NFSv4; monitor NFS server health and network latency; use automount (autofs) to unmount idle NFS shares automatically.

QUESTION 17LinuxHard

Your production server disk is at 100% according to df -h but du -sh / shows only 60% usage. What causes this and how do you recover?

#
Reveal answer guidance

Discrepancy means a process holds open file descriptors to deleted files. When a file is deleted while open, du no longer counts it (no directory entry), but df still includes it because the inode is referenced. Find the offenders: lsof | grep "(deleted)" | awk '{print $2, $NF}' to list PID and file path of deleted files. Then check ls -la /proc/<PID>/fd/ | grep deleted for exact sizes. Recovery options: (1) Kill the process holding the file — the kernel releases the space: kill <PID>. (2) If the process cannot be restarted, truncate the file: : > /proc/<PID>/fd/<FD_NUM> — this empties the file content in-place, releasing disk space without interrupting the process. (3) For log files that grow continuously, configure logrotate with copytruncate (copies log then truncates) instead of create (creates new inode) to avoid this scenario. Preventive monitoring: alert on disk_usage > 80% and track df -i for inode exhaustion separately.

QUESTION 18LinuxHard

A process uses 99% CPU but the OOM killer keeps killing other processes instead. Explain the OOM scoring algorithm and how to target the correct process.

#
Reveal answer guidance

The OOM killer uses oom_score = total_rss + total_swap + (total_pages * page_table_factor) / oom_score_adj. The key factor: oom_score_adj (range -1000 to 1000). By default, all processes have oom_score_adj=0. A CPU-intensive process that uses little memory (most CPU time is spent in tight loops on cached/pinned data) has a low oom_score because RSS and swap are low. The OOM killer targets the process with the highest oom_score, which is typically a memory-intensive process (database, cache) that happens to be using lots of RAM. To protect memory-heavy but critical processes, set echo -500 > /proc/<critical_pid>/oom_score_adj (less likely killed). To target the CPU-hungry process: echo 500 > /proc/<cpu_hog_pid>/oom_score_adj which increases its score. But the best fix is addressing the actual problem — either reduce memory pressure or limit the CPU hog. Use systemd-run --scope -p MemoryMax=<limit> -p CPUQuota=<limit> <command> to constrain the process. For production: never rely on OOM killer — set appropriate cgroup limits (memory.max in cgroup v2) per service so the kernel OOMs only within the cgroup, not globally.

QUESTION 19LinuxHard

Linux incident: A Linux server shows 100% CPU but low application throughput. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: top, mpstat, pidstat, perf top, run queue length, and application profiling. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Check user/system/iowait/steal time, run queue, hot threads, and recent deployments. Tune the bottleneck, not just instance size.

QUESTION 20LinuxHard

Linux architecture scenario: You need to separate CPU-bound work from lock contention and IO wait. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Check user/system/iowait/steal time, run queue, hot threads, and recent deployments. Tune the bottleneck, not just instance size.

QUESTION 21LinuxMedium

Linux security scenario: A quick restart hides the evidence needed for root cause. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Check user/system/iowait/steal time, run queue, hot threads, and recent deployments. Tune the bottleneck, not just instance size.

QUESTION 22LinuxHard

Linux release scenario: You are changing CPU limits and process placement. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Check user/system/iowait/steal time, run queue, hot threads, and recent deployments. Tune the bottleneck, not just instance size.

QUESTION 23LinuxMedium

Linux reliability/cost scenario: Peak-hour CPU alarms are firing every day. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Check user/system/iowait/steal time, run queue, hot threads, and recent deployments. Tune the bottleneck, not just instance size.

QUESTION 24LinuxHard

Linux incident: The kernel OOM killer terminates critical processes at random. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: dmesg, journalctl, /proc/meminfo, smem, ps aux --sort -rss, and cgroup memory stats. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Identify RSS growth, cache behavior, cgroup limits, and OOM scores. Add memory alerts, fix leaks, and protect critical services with sane limits.

QUESTION 25LinuxHard

Linux architecture scenario: You need predictable memory behavior and early warning. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Identify RSS growth, cache behavior, cgroup limits, and OOM scores. Add memory alerts, fix leaks, and protect critical services with sane limits.

QUESTION 26LinuxMedium

Linux security scenario: Swap hides memory leaks until latency collapses. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Identify RSS growth, cache behavior, cgroup limits, and OOM scores. Add memory alerts, fix leaks, and protect critical services with sane limits.

QUESTION 27LinuxHard

Linux release scenario: You are introducing systemd memory controls. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Identify RSS growth, cache behavior, cgroup limits, and OOM scores. Add memory alerts, fix leaks, and protect critical services with sane limits.

QUESTION 28LinuxMedium

Linux reliability/cost scenario: Memory upgrades are increasing cost without fixing leaks. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Identify RSS growth, cache behavior, cgroup limits, and OOM scores. Add memory alerts, fix leaks, and protect critical services with sane limits.

QUESTION 29LinuxHard

Linux incident: The root filesystem fills up and services fail to write logs or temp files. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: df -h, du -xhd1, lsof +L1, journal size, inode usage, and logrotate status. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Check both bytes and inodes, find deleted-open files, rotate logs, separate data volumes, and alert before critical filesystems exceed safe thresholds.

QUESTION 30LinuxHard

Linux architecture scenario: You need disk isolation for logs, data, and system partitions. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Check both bytes and inodes, find deleted-open files, rotate logs, separate data volumes, and alert before critical filesystems exceed safe thresholds.

QUESTION 31LinuxMedium

Linux security scenario: Deleting files that are still held open does not free space. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Check both bytes and inodes, find deleted-open files, rotate logs, separate data volumes, and alert before critical filesystems exceed safe thresholds.

QUESTION 32LinuxHard

Linux release scenario: You are moving logs and data to separate mount points. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Check both bytes and inodes, find deleted-open files, rotate logs, separate data volumes, and alert before critical filesystems exceed safe thresholds.

QUESTION 33LinuxMedium

Linux reliability/cost scenario: Emergency disk cleanup is becoming routine. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Check both bytes and inodes, find deleted-open files, rotate logs, separate data volumes, and alert before critical filesystems exceed safe thresholds.

QUESTION 34LinuxHard

Linux incident: Only one service path has high latency from a Linux host. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: curl -w, dig, ss -ti, mtr, tcpdump, MTU checks, and conntrack stats. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Break latency into lookup, connect, TLS, first byte, and transfer. Validate route, MTU, retransmits, and connection reuse before blaming the app.

QUESTION 35LinuxHard

Linux architecture scenario: You need to isolate DNS, TCP, TLS, routing, and application latency. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Break latency into lookup, connect, TLS, first byte, and transfer. Validate route, MTU, retransmits, and connection reuse before blaming the app.

QUESTION 36LinuxMedium

Linux security scenario: Packet captures may expose sensitive payloads. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Break latency into lookup, connect, TLS, first byte, and transfer. Validate route, MTU, retransmits, and connection reuse before blaming the app.

QUESTION 37LinuxHard

Linux release scenario: You are changing MTU or routing in production. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Break latency into lookup, connect, TLS, first byte, and transfer. Validate route, MTU, retransmits, and connection reuse before blaming the app.

QUESTION 38LinuxMedium

Linux reliability/cost scenario: Retries are amplifying downstream load. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Break latency into lookup, connect, TLS, first byte, and transfer. Validate route, MTU, retransmits, and connection reuse before blaming the app.

QUESTION 39LinuxHard

Linux incident: A systemd service repeatedly restarts but manual command execution works. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: systemctl status, journalctl -u, systemd-analyze verify, unit environment, and exit codes. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Compare interactive shell versus unit environment, working directory, user, limits, dependencies, and sandboxing. Add explicit env files and meaningful restart policies.

QUESTION 40LinuxHard

Linux architecture scenario: You need reliable service units with correct environment and dependencies. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Compare interactive shell versus unit environment, working directory, user, limits, dependencies, and sandboxing. Add explicit env files and meaningful restart policies.

QUESTION 41LinuxMedium

Linux security scenario: Running services as root masks permission issues. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Compare interactive shell versus unit environment, working directory, user, limits, dependencies, and sandboxing. Add explicit env files and meaningful restart policies.

QUESTION 42LinuxHard

Linux release scenario: You are adding hardening options to systemd units. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Compare interactive shell versus unit environment, working directory, user, limits, dependencies, and sandboxing. Add explicit env files and meaningful restart policies.

QUESTION 43LinuxMedium

Linux reliability/cost scenario: Restart loops page on-call with no useful context. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Compare interactive shell versus unit environment, working directory, user, limits, dependencies, and sandboxing. Add explicit env files and meaningful restart policies.

QUESTION 44LinuxHard

Linux incident: A service fails with too many open files during traffic spikes. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: lsof -p, /proc/<pid>/limits, ss -s, systemd limits, and application connection pool metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Find whether descriptors are sockets, files, or leaks. Tune per-service limits, pool sizes, keepalive, and alert on descriptor growth.

QUESTION 45LinuxHard

Linux architecture scenario: You need safe connection scaling. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Find whether descriptors are sockets, files, or leaks. Tune per-service limits, pool sizes, keepalive, and alert on descriptor growth.

QUESTION 46LinuxMedium

Linux security scenario: Raising limits globally hides leaks and increases blast radius. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Find whether descriptors are sockets, files, or leaks. Tune per-service limits, pool sizes, keepalive, and alert on descriptor growth.

QUESTION 47LinuxHard

Linux release scenario: You are changing LimitNOFILE and kernel limits. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Find whether descriptors are sockets, files, or leaks. Tune per-service limits, pool sizes, keepalive, and alert on descriptor growth.

QUESTION 48LinuxMedium

Linux reliability/cost scenario: Connection errors appear during every load test. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Find whether descriptors are sockets, files, or leaks. Tune per-service limits, pool sizes, keepalive, and alert on descriptor growth.

QUESTION 49LinuxHard

Linux incident: High connection churn causes port exhaustion. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: ss -s, ephemeral port range, TIME_WAIT counts, conntrack usage, and retry metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Tune connection reuse, keepalive, pool sizes, and ephemeral port range carefully. Prefer application pooling before aggressive kernel changes.

QUESTION 50LinuxHard

Linux architecture scenario: You need stable outbound connectivity for high-throughput clients. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Tune connection reuse, keepalive, pool sizes, and ephemeral port range carefully. Prefer application pooling before aggressive kernel changes.

QUESTION 51LinuxMedium

Linux security scenario: Unsafe sysctl changes can break unrelated workloads. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Tune connection reuse, keepalive, pool sizes, and ephemeral port range carefully. Prefer application pooling before aggressive kernel changes.

QUESTION 52LinuxHard

Linux release scenario: You are changing TCP sysctls on production hosts. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Tune connection reuse, keepalive, pool sizes, and ephemeral port range carefully. Prefer application pooling before aggressive kernel changes.

QUESTION 53LinuxMedium

Linux reliability/cost scenario: NAT and client hosts are saturating ephemeral ports. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Tune connection reuse, keepalive, pool sizes, and ephemeral port range carefully. Prefer application pooling before aggressive kernel changes.

QUESTION 54LinuxHard

Linux incident: A deploy fails because the service user cannot read a new config file. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: namei -l, stat, ACL checks, SELinux/AppArmor status, and service user identity. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Check every parent directory permission, ownership, ACLs, and MAC policy. Use dedicated service users and deployment automation that sets ownership deterministically.

QUESTION 55LinuxHard

Linux architecture scenario: You need least-privilege filesystem access that survives deploys. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Check every parent directory permission, ownership, ACLs, and MAC policy. Use dedicated service users and deployment automation that sets ownership deterministically.

QUESTION 56LinuxMedium

Linux security scenario: Using chmod 777 creates a security incident. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Check every parent directory permission, ownership, ACLs, and MAC policy. Use dedicated service users and deployment automation that sets ownership deterministically.

QUESTION 57LinuxHard

Linux release scenario: You are standardizing service users and file ownership. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Check every parent directory permission, ownership, ACLs, and MAC policy. Use dedicated service users and deployment automation that sets ownership deterministically.

QUESTION 58LinuxMedium

Linux reliability/cost scenario: Permission fixes are manual and inconsistent. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Check every parent directory permission, ownership, ACLs, and MAC policy. Use dedicated service users and deployment automation that sets ownership deterministically.

QUESTION 59LinuxHard

Linux incident: A process gets permission denied even though Unix permissions look correct. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: ausearch, audit2why, AppArmor logs, context labels, and policy status. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Confirm MAC denial, fix labels or profiles narrowly, test in permissive/report mode, and avoid blanket disablement.

QUESTION 60LinuxHard

Linux architecture scenario: You need mandatory access control without disabling it globally. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Confirm MAC denial, fix labels or profiles narrowly, test in permissive/report mode, and avoid blanket disablement.

QUESTION 61LinuxMedium

Linux security scenario: Turning enforcement off hides a real policy violation. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Confirm MAC denial, fix labels or profiles narrowly, test in permissive/report mode, and avoid blanket disablement.

QUESTION 62LinuxHard

Linux release scenario: You are moving from permissive to enforcing mode. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Confirm MAC denial, fix labels or profiles narrowly, test in permissive/report mode, and avoid blanket disablement.

QUESTION 63LinuxMedium

Linux reliability/cost scenario: Security requires host hardening before production approval. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Confirm MAC denial, fix labels or profiles narrowly, test in permissive/report mode, and avoid blanket disablement.

QUESTION 64LinuxHard

Linux incident: TLS, logs, and distributed traces disagree because host clocks drift. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: timedatectl, chronyc tracking, NTP source checks, and drift metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use a reliable time source, monitor offset, avoid manual time jumps, and make alerting sensitive to clock drift.

QUESTION 65LinuxHard

Linux architecture scenario: You need reliable time across fleets. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use a reliable time source, monitor offset, avoid manual time jumps, and make alerting sensitive to clock drift.

QUESTION 66LinuxMedium

Linux security scenario: Manual time changes can break databases and distributed locks. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use a reliable time source, monitor offset, avoid manual time jumps, and make alerting sensitive to clock drift.

QUESTION 67LinuxHard

Linux release scenario: You are standardizing chrony or systemd-timesyncd. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use a reliable time source, monitor offset, avoid manual time jumps, and make alerting sensitive to clock drift.

QUESTION 68LinuxMedium

Linux reliability/cost scenario: Debugging incidents is slow because timestamps cannot be trusted. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use a reliable time source, monitor offset, avoid manual time jumps, and make alerting sensitive to clock drift.

QUESTION 69LinuxHard

Linux incident: Load average is high but CPU is not fully used. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: uptime, top, vmstat, iostat, blocked process stacks, and run queue analysis. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Load includes runnable and uninterruptible tasks. Check IO wait, blocked tasks, storage latency, and locks before buying more CPU.

QUESTION 70LinuxHard

Linux architecture scenario: You need to explain load average correctly in an interview. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Load includes runnable and uninterruptible tasks. Check IO wait, blocked tasks, storage latency, and locks before buying more CPU.

QUESTION 71LinuxMedium

Linux security scenario: Scaling CPU does not fix IO or uninterruptible sleep. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Load includes runnable and uninterruptible tasks. Check IO wait, blocked tasks, storage latency, and locks before buying more CPU.

QUESTION 72LinuxHard

Linux release scenario: You are changing storage or process concurrency. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Load includes runnable and uninterruptible tasks. Check IO wait, blocked tasks, storage latency, and locks before buying more CPU.

QUESTION 73LinuxMedium

Linux reliability/cost scenario: Executives see high load and ask for bigger instances. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Load includes runnable and uninterruptible tasks. Check IO wait, blocked tasks, storage latency, and locks before buying more CPU.

QUESTION 74LinuxHard

Linux incident: Database latency spikes line up with disk await spikes. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: iostat -xz, pidstat -d, filesystem errors, queue depth, and database IO metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Check await, utilization, queue depth, fsync rate, and noisy neighbors. Tune application write patterns and storage class together.

QUESTION 75LinuxHard

Linux architecture scenario: You need to separate filesystem, block device, and application causes. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Check await, utilization, queue depth, fsync rate, and noisy neighbors. Tune application write patterns and storage class together.

QUESTION 76LinuxMedium

Linux security scenario: Filesystem remounts can cause downtime. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Check await, utilization, queue depth, fsync rate, and noisy neighbors. Tune application write patterns and storage class together.

QUESTION 77LinuxHard

Linux release scenario: You are changing disk class or mount options. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Check await, utilization, queue depth, fsync rate, and noisy neighbors. Tune application write patterns and storage class together.

QUESTION 78LinuxMedium

Linux reliability/cost scenario: IOPS costs are high but latency remains unstable. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Check await, utilization, queue depth, fsync rate, and noisy neighbors. Tune application write patterns and storage class together.

QUESTION 79LinuxHard

Linux incident: Applications pause for seconds before connecting to dependencies. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: resolvectl, dig +trace, /etc/resolv.conf, packet captures, and resolver logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Check resolver order, search domains, timeout/retry values, negative caching, and application DNS cache behavior.

QUESTION 80LinuxHard

Linux architecture scenario: You need fast and reliable resolver behavior. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Check resolver order, search domains, timeout/retry values, negative caching, and application DNS cache behavior.

QUESTION 81LinuxMedium

Linux security scenario: Search domains leak internal names to upstream DNS. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Check resolver order, search domains, timeout/retry values, negative caching, and application DNS cache behavior.

QUESTION 82LinuxHard

Linux release scenario: You are changing resolver config fleet-wide. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Check resolver order, search domains, timeout/retry values, negative caching, and application DNS cache behavior.

QUESTION 83LinuxMedium

Linux reliability/cost scenario: Connection latency is dominated by DNS. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Check resolver order, search domains, timeout/retry values, negative caching, and application DNS cache behavior.

QUESTION 84LinuxHard

Linux incident: Zombie or orphaned processes accumulate after deployments. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: ps -eo, parent PID inspection, systemd cgroup view, and signal handling tests. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Ensure PID 1 reaps children, services handle SIGTERM, and supervisors do not double-fork unexpectedly. Alert on process count growth.

QUESTION 85LinuxHard

Linux architecture scenario: You need clean process lifecycle management. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Ensure PID 1 reaps children, services handle SIGTERM, and supervisors do not double-fork unexpectedly. Alert on process count growth.

QUESTION 86LinuxMedium

Linux security scenario: PID exhaustion can take down the host. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Ensure PID 1 reaps children, services handle SIGTERM, and supervisors do not double-fork unexpectedly. Alert on process count growth.

QUESTION 87LinuxHard

Linux release scenario: You are adding an init process or fixing signal handling. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Ensure PID 1 reaps children, services handle SIGTERM, and supervisors do not double-fork unexpectedly. Alert on process count growth.

QUESTION 88LinuxMedium

Linux reliability/cost scenario: Hosts require periodic reboot to stay healthy. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Ensure PID 1 reaps children, services handle SIGTERM, and supervisors do not double-fork unexpectedly. Alert on process count growth.

QUESTION 89LinuxHard

Linux incident: A security patch restarts a dependency and causes application errors. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: package changelogs, service restart list, canary host metrics, and maintenance window logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Patch canaries first, know which services restart, drain traffic, snapshot critical hosts, and separate emergency security fixes from routine upgrades.

QUESTION 90LinuxHard

Linux architecture scenario: You need safe Linux patching for production fleets. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Patch canaries first, know which services restart, drain traffic, snapshot critical hosts, and separate emergency security fixes from routine upgrades.

QUESTION 91LinuxMedium

Linux security scenario: Unattended upgrades change runtime behavior without app owners knowing. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Patch canaries first, know which services restart, drain traffic, snapshot critical hosts, and separate emergency security fixes from routine upgrades.

QUESTION 92LinuxHard

Linux release scenario: You are introducing rolling OS patch windows. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Patch canaries first, know which services restart, drain traffic, snapshot critical hosts, and separate emergency security fixes from routine upgrades.

QUESTION 93LinuxMedium

Linux reliability/cost scenario: Vulnerability SLAs require faster patch rollout. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Patch canaries first, know which services restart, drain traffic, snapshot critical hosts, and separate emergency security fixes from routine upgrades.

QUESTION 94LinuxHard

Linux incident: A port listens locally but remote clients cannot connect. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: ss -ltnp, iptables/nft rules, cloud firewall rules, route checks, and packet counters. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Verify listener, local firewall, route, upstream firewall, and return path. Open the narrowest source, destination, and port required.

QUESTION 95LinuxHard

Linux architecture scenario: You need layered host and network firewall troubleshooting. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Verify listener, local firewall, route, upstream firewall, and return path. Open the narrowest source, destination, and port required.

QUESTION 96LinuxMedium

Linux security scenario: Opening broad CIDRs exposes admin services. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Verify listener, local firewall, route, upstream firewall, and return path. Open the narrowest source, destination, and port required.

QUESTION 97LinuxHard

Linux release scenario: You are migrating iptables rules to nftables. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Verify listener, local firewall, route, upstream firewall, and return path. Open the narrowest source, destination, and port required.

QUESTION 98LinuxMedium

Linux reliability/cost scenario: Network tickets bounce between teams. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Verify listener, local firewall, route, upstream firewall, and return path. Open the narrowest source, destination, and port required.

QUESTION 99LinuxHard

Linux incident: Backups exist but restore has never been tested. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: backup logs, checksum validation, restore test output, and application consistency checks. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Test restores regularly, capture RTO/RPO, validate application-level consistency, encrypt backups, and restrict delete permissions.

QUESTION 100LinuxHard

Linux architecture scenario: You need provable recovery for Linux-hosted data. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Test restores regularly, capture RTO/RPO, validate application-level consistency, encrypt backups, and restrict delete permissions.

QUESTION 101LinuxMedium

Linux security scenario: Backups may contain corrupted or incomplete files. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Test restores regularly, capture RTO/RPO, validate application-level consistency, encrypt backups, and restrict delete permissions.

QUESTION 102LinuxHard

Linux release scenario: You are adding automated restore drills. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Test restores regularly, capture RTO/RPO, validate application-level consistency, encrypt backups, and restrict delete permissions.

QUESTION 103LinuxMedium

Linux reliability/cost scenario: Audit requires recovery time and recovery point evidence. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Test restores regularly, capture RTO/RPO, validate application-level consistency, encrypt backups, and restrict delete permissions.

QUESTION 104LinuxHard

Linux incident: An incident has too many logs and no clear timeline. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: journalctl --since, app logs, auth logs, kernel logs, and correlation IDs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Build a timeline by host, service, deployment, and user impact. Centralize logs, preserve evidence, and redact sensitive fields before broad sharing.

QUESTION 105LinuxHard

Linux architecture scenario: You need fast Linux incident reconstruction. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Build a timeline by host, service, deployment, and user impact. Centralize logs, preserve evidence, and redact sensitive fields before broad sharing.

QUESTION 106LinuxMedium

Linux security scenario: Logs may include secrets or sensitive request data. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Build a timeline by host, service, deployment, and user impact. Centralize logs, preserve evidence, and redact sensitive fields before broad sharing.

QUESTION 107LinuxHard

Linux release scenario: You are standardizing structured logs and retention. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Build a timeline by host, service, deployment, and user impact. Centralize logs, preserve evidence, and redact sensitive fields before broad sharing.

QUESTION 108LinuxMedium

Linux reliability/cost scenario: Mean time to resolution is high. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Build a timeline by host, service, deployment, and user impact. Centralize logs, preserve evidence, and redact sensitive fields before broad sharing.

QUESTION 109LinuxHard

Linux incident: A Linux host fails to boot after a configuration change. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: serial console, rescue mode, journalctl -b -1, fstab validation, and disk checks. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use console access, boot rescue, validate fstab and kernel args, snapshot before repair, and automate config validation before rollout.

QUESTION 110LinuxHard

Linux architecture scenario: You need a recovery path without data loss. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use console access, boot rescue, validate fstab and kernel args, snapshot before repair, and automate config validation before rollout.

QUESTION 111LinuxMedium

Linux security scenario: Repeated repair attempts can damage disks or overwrite evidence. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use console access, boot rescue, validate fstab and kernel args, snapshot before repair, and automate config validation before rollout.

QUESTION 112LinuxHard

Linux release scenario: You are changing bootloader, fstab, or kernel settings. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use console access, boot rescue, validate fstab and kernel args, snapshot before repair, and automate config validation before rollout.

QUESTION 113LinuxMedium

Linux reliability/cost scenario: Manual host recovery delays service restoration. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use console access, boot rescue, validate fstab and kernel args, snapshot before repair, and automate config validation before rollout.

QUESTION 114LinuxHard

Linux incident: A contractor still has SSH access after leaving the project. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: last, auth logs, authorized keys audit, sudoers review, and identity provider logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use individual identities, centralized SSH certificates or SSO, short-lived access, audited sudo, and automated offboarding.

QUESTION 115LinuxHard

Linux architecture scenario: You need auditable Linux access management. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use individual identities, centralized SSH certificates or SSO, short-lived access, audited sudo, and automated offboarding.

QUESTION 116LinuxMedium

Linux security scenario: Shared users prevent accountability. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use individual identities, centralized SSH certificates or SSO, short-lived access, audited sudo, and automated offboarding.

QUESTION 117LinuxHard

Linux release scenario: You are moving from static SSH keys to centralized access. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use individual identities, centralized SSH certificates or SSO, short-lived access, audited sudo, and automated offboarding.

QUESTION 118LinuxMedium

Linux reliability/cost scenario: Compliance requires proof of least privilege. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use individual identities, centralized SSH certificates or SSO, short-lived access, audited sudo, and automated offboarding.

CONTINUE PRACTICING

Try another perspective.