PRACTICE TRACK / 18 QUESTIONS

Python
Think it through.

Concurrency, performance, async, and SRE tooling.

Choose a question, explain your approach, then reveal the supplied answer. Difficulty labels come from the existing question library.

18 questions

Answers stay closed until you choose to reveal them.

QUESTION 01PythonEasy

What is the difference between os.environ.get("KEY") and os.environ["KEY"], and which pattern should SRE automation scripts prefer?

#
Reveal answer guidance

os.environ["KEY"] raises KeyError if the variable is absent, crashing the script. os.environ.get("KEY") returns None (or a supplied default). SRE scripts should prefer os.environ.get("KEY", "default") at startup and then validate explicitly: if not value: raise SystemExit("KEY env var required"). Better still, use pydantic-settings or python-dotenv with a typed settings model — you get type coercion, validation, and a single crash-fast entry point rather than scattered KeyErrors deep in business logic.

QUESTION 02PythonMedium

What is the difference between asyncio.gather() and asyncio.TaskGroup (Python 3.11+), and when does each matter for SRE scripts that fan out to multiple API endpoints?

#
Reveal answer guidance

asyncio.gather() has a subtle hazard: if one coroutine raises and return_exceptions=False (default), other tasks are NOT cancelled — they keep running, leaking resources and making errors hard to trace. asyncio.TaskGroup uses structured concurrency: if any task raises, all sibling tasks are immediately cancelled and the exception propagates cleanly out of the async with block. For SRE fan-out (querying 50 Prometheus endpoints simultaneously), prefer TaskGroup because: (1) partial failures cancel the rest, preventing stale data aggregation; (2) the stack trace is much cleaner; (3) it integrates with asyncio.timeout() for per-group deadlines. Use gather(return_exceptions=True) only when you explicitly want to collect all errors without cancellation — e.g., health-check polling where partial success is meaningful.

QUESTION 03PythonMedium

What is the GIL, and how does it affect an SRE observability agent that needs to process metrics from 200 targets concurrently?

#
Reveal answer guidance

The GIL (Global Interpreter Lock) allows only one thread to execute Python bytecode at a time. For CPU-bound work, threads give no parallelism — use multiprocessing. For I/O-bound work (HTTP scrapes, DNS lookups), the GIL is released during blocking syscalls, so threading or asyncio both achieve true concurrency. For 200 Prometheus scrape targets: (1) Best: asyncio with aiohttp — single thread, no GIL contention, scales to thousands of targets with minimal memory; (2) Good: ThreadPoolExecutor(max_workers=50) with requests — GIL released during HTTP I/O, but thread overhead is higher; (3) Avoid: ProcessPoolExecutor for pure scraping — cross-process pickle serialization adds latency and memory. Python 3.13 introduces a no-GIL build (PEP 703) but it's experimental — don't rely on it in production SRE tooling yet.

QUESTION 04PythonHard

You need to write a Python-based SRE tool that parses 10 GB of compressed Nginx access logs (*.gz) in under 60 seconds on a 16-core machine. Walk through the complete architecture, including the concurrency model, memory management, and parsing strategy.

#
Reveal answer guidance

Architecture: (1) Discovery: glob.glob("/var/log/nginx/*.gz") — sort by size descending for better load balancing. (2) Parallelism: multiprocessing.Pool(processes=os.cpu_count()) — GIL is irrelevant for CPU-bound gzip decompression + regex; each process is truly independent. (3) Memory: never load an entire file into RAM. Use gzip.open(f, "rt", encoding="utf-8", errors="replace") as a streaming iterator. Each worker processes its assigned file line-by-line, maintaining only a local collections.Counter. (4) Parsing: pre-compile the regex once per process — compiled regex is ~5x faster than string methods for complex patterns. (5) IPC: each worker returns a small Counter dict; main process merges with sum(counters, Counter()). (6) Chunk assignment: assign whole files per worker (not line ranges) to exploit OS page cache locality. (7) Bottleneck check: avoid datetime.strptime — use manual string slicing for timestamp fields. Expected: ~200 MB/s compressed per core × 16 cores easily processes 10 GB in under 60s.

QUESTION 05PythonHard

A Python-based Kubernetes operator using the kopf framework is experiencing event storms: a single ConfigMap change triggers 10,000 handler invocations in 30 seconds, OOMing the operator pod. Diagnose the root cause and redesign the handler to be storm-resilient.

#
Reveal answer guidance

Root cause candidates: (1) Update loop — the handler modifies the ConfigMap it's watching (e.g., to add a last-synced annotation), triggering another watch event creating infinite recursion. (2) Missing ownership filter — the handler watches all ConfigMaps but the operator owns thousands; a cluster-wide change triggers all handlers simultaneously. (3) kopf retry storm — handler raises an exception, kopf schedules immediate retry, which also fails, filling the retry queue. (4) No debounce — kopf's default config doesn't throttle rapid updates to the same object. Fixes: (1) Use patch.status["last-synced"] = ... via kopf's built-in patch object — kopf suppresses events caused by its own patches using the kopf.dev/last-handled-configuration annotation. (2) Ownership filter: @kopf.on.update("", "v1", "configmaps", labels={"managed-by": "my-operator"}). (3) Exponential backoff: set @kopf.on.update(..., backoff=60) and retries=3. (4) Debounce: @kopf.on.update(..., initial_delay=5.0) batches rapid updates. (5) Worker pool limit: --threads=4 CLI flag.

QUESTION 06PythonHard

A decorator applied to a coroutine function does not preserve the inspect.iscoroutinefunction check downstream. Write a decorator that properly wraps both sync and async functions without breaking awaitability or type introspection.

#
Reveal answer guidance

When a decorator returns a regular function for an async def, inspect.iscoroutinefunction returns False because the wrapper is a synchronous function. The fix: use @functools.wraps on an async def wrapper for async targets, and a sync wrapper for sync targets. Detect the target at decoration time: if asyncio.iscoroutinefunction(func), return an async wrapper; otherwise return a sync wrapper. Use functools.wraps(func) on the appropriate wrapper to copy __name__, __qualname__, __doc__, __module__, and __dict__. However, @functools.wraps does not set __wrapped__ on coroutine functions correctly for inspect.unwrap. For bulletproof introspection, set __signature__ from inspect.signature(func). Alternatively, use a class-based decorator with __call__ returning await if the instance detects a coroutine — but this is fragile. The Python 3.10+ inspect.markcoroutinefunction can mark the wrapper, but the cleanest pattern: use @functools.singledispatch on the decorator, dispatching on isinstance(func, (types.FunctionType, types.MethodType)) and returning the appropriately typed wrapper.

QUESTION 07PythonHard

What is the difference between @dataclass and typing.NamedTuple beyond mutability? When would using a NamedTuple cause subtle bugs that a dataclass would avoid?

#
Reveal answer guidance

NamedTuple extends tuple — instances are immutable and hashable (if all fields are hashable), and they inherit __lt__, __eq__, etc. from tuple, so comparison is field-by-field (like tuple ordering). This causes a bug: NamedTuple instances compare equal to plain tuples of the same length and values — MyNamedTuple(x=1, y=2) == (1, 2) returns True, which violates type expectations. @dataclass generates __eq__ that checks type identity (isinstance check, unless eq=False). NamedTuple also serializes as a tuple via pickle, not as a dict — so JSON serialization is not automatic. @dataclass supports __slots__ (Python 3.10+ via slots=True), default factories, __post_init__ hooks, and frozen=True for immutability. NamedTuple supports only simple field defaults (no mutable defaults without workaround via None + __new__). The NamedTuple is better for: lightweight immutable records where tuple-like unpacking is desired (x, y = point), and when inheriting from tuple is an explicit design choice (e.g., compatibility with tuple-based APIs). For all other data containers, @dataclass is superior.

QUESTION 08PythonMedium

Explain how contextlib.contextmanager works internally. What happens if the generator yields multiple times or does not yield at all? Can you use it with async context managers?

#
Reveal answer guidance

contextlib.contextmanager wraps a generator function into a context manager. When __enter__ is called, it creates a generator iterator and calls next(gen), which executes up to the yield. The yielded value is returned from __enter__. When __exit__ is called, it calls next(gen) again — this time the generator runs the cleanup code and should raise StopIteration. If the generator yields multiple times, __exit__ calls next(gen) repeatedly, which is confusing — it will raise StopIteration on the second next() call, but the first next() already executed cleanup, causing double-free bugs. If the generator never yields, __enter__ raises RuntimeError("generator didn't yield"). If an exception occurs in the with block, it's passed to gen.throw(exc_type, exc_val, exc_tb) inside __exit__. For async: @contextlib.asynccontextmanager works the same but with async def and __aenter__/__aexit__. The async version uses anext(gen) and gen.athrow() instead.

QUESTION 09PythonHard

How does functools.lru_cache implement its eviction policy? What happens to the internal CacheInfo when the cache is called from multiple threads without thread_safe (Python 3.9+)?

#
Reveal answer guidance

lru_cache uses a doubly-linked list combined with a hash table. Each cache entry has prev and next pointers forming an ordered list by access time (MRU at head, LRU at tail). On a cache hit, the entry is unlinked and moved to the head (O(1)). On a miss when full, the tail entry is evicted (O(1)). The hash table maps keys to linked-list nodes. Without thread_safe=True (added in Python 3.9), the C implementation (_functools.c) does not hold a lock during mutation — two threads can hit a miss simultaneously and both compute the value, with a race on linked-list pointers leading to corruption or infinite loops. The CacheInfo counters (hits, misses, currsize, maxsize) use non-atomic increments, so info() may show inconsistent values under concurrent access. With thread_safe=True, a PyThread_type_lock protects all mutations and queries. The lock is a single global lock per cache, so high-concurrency workloads may still contend. For concurrent workloads with cache-thrashing, prefer a specialized library like cachetools.TTLCache with its own locking strategy.

QUESTION 10PythonMedium

What is the difference between Protocol and ABC in typing? When would a Protocol be strictly necessary that ABC cannot replace?

#
Reveal answer guidance

ABC enforces explicit inheritance — a class must inherit from the ABC and implement the abstract methods for isinstance() to return True. This is nominal subtyping. Protocol in Python 3.8+ (via typing_extensions for earlier versions) enables structural subtyping (duck typing at the type-checker level): any class that has the required methods and attributes is compatible with the Protocol without needing to inherit from it. Protocol is necessary when working with third-party libraries whose classes you cannot modify but whose interfaces match expected signatures. For example, if a function accepts any "openable" object with a .read() method, a Protocol class with def read(self) -> bytes allows mypy to accept io.BytesIO, tempfile.SpooledTemporaryFile, and even custom file-like objects without inheritance. An ABC would require each class to explicitly subclass it. At runtime, both ABC and Protocol support isinstance (if @runtime_checkable decorates the Protocol), but @runtime_checkable only checks for attribute presence, not method signatures. For static analysis only, Protocol is preferred for flexibility.

QUESTION 11PythonHard

You write a metaclass that intercepts __init_subclass__ calls. The metaclass must track all subclasses while ensuring that __init_subclass__ in the child class is not silently overridden. How do you resolve this collision?

#
Reveal answer guidance

When both a metaclass and __init_subclass__ are used, the metaclass's __init_subclass__ (if defined) is called first, then the class's __init_subclass__. But if the child class defines its own __init_subclass__, it shadows the metaclass's hook. To ensure the metaclass always tracks subclasses regardless, override __new__ in the metaclass instead: capture the new class after super().__new__() is called, then register it. This bypasses __init_subclass__ entirely. Example: class Meta(type): def __new__(mcs, name, bases, ns, **kwargs): cls = super().__new__(mcs, name, bases, ns); Meta._registry[name] = cls; return cls. For the __init_subclass__ collision: the metaclass can store its own callback in the class namespace before type.__new__ is invoked, then have __init_subclass__ call super().__init_subclass__() explicitly. Alternatively, use __set_name__ on a descriptor to wire up hooks. The cleanest design: avoid overriding __init_subclass__ in the metaclass — track subclasses in __init__ instead, which fires after the class dict is created.

QUESTION 12PythonHard

You have a generator that yields 10 million rows from a database. The consumer breaks after 5,000 rows and never reads the rest. What happens to the generator's resources (connection, cursor) and how do you ensure proper cleanup?

#
Reveal answer guidance

When the consumer stops iterating, the generator function is suspended at the last yield. The generator's __del__ method is called when reference count drops to zero. If the generator has a finally block, it runs during throw(GeneratorExit) inside __del__. However, if an exception traceback references the generator frame, cyclic garbage is involved, and cleanup is delayed until the next GC cycle. This means the database connection may remain open. To ensure prompt cleanup, wrap the generator in a context manager: @contextmanager closes the connection on __exit__. Or use GeneratorExit handling: the generator catches GeneratorExit (raised by .close()) and runs cleanup. The consumer should call .close() on the generator when done early, which raises GeneratorExit at the yield point. However, most consumers (for loops) do not call .close() unless they break — and break implicitly calls .close(). The real danger is when the consumer raises an exception: the generator's finally runs, but if the finally also accesses the connection, it may fail if the connection was already in an error state. The safest pattern: use contextlib.closing(generator) or wrap in a connection context manager that lives outside the generator.

QUESTION 13PythonMedium

Explain the difference between concurrent.futures.ThreadPoolExecutor and asyncio for I/O-bound tasks. When would you choose ThreadPoolExecutor over async/await despite the GIL?

#
Reveal answer guidance

Both handle I/O-bound tasks, but they differ fundamentally: ThreadPoolExecutor uses OS threads waiting on blocking I/O (each thread occupies a C stack ~8MB), while asyncio uses a single thread with cooperative multitasking — tasks yield at await points. ThreadPoolExecutor is simpler for existing blocking libraries (requests, psycopg2, boto3 without async support) because you don't need to rewrite code. asyncio requires libraries with async interfaces (aiohttp, asyncpg, aiobotocore). Despite the GIL, ThreadPoolExecutor is effective for I/O because the GIL is released during blocking syscalls (read/write/connect). Choose ThreadPoolExecutor when: (1) your dependencies are blocking and cannot be replaced, (2) you need true parallelism for CPU-light work mixed with I/O, (3) you have a small number of long-running I/O tasks (e.g., downloading 10 large files). Choose asyncio when: (1) you have thousands of concurrent connections (WebSocket servers, proxies), (2) you need fine-grained control over task scheduling, (3) you want lower memory overhead — asyncio tasks are ~1KB vs threads ~8MB stack. In practice, hybrid patterns exist: loop.run_in_executor() submits blocking calls to a thread pool while the event loop runs async code.

QUESTION 14PythonMedium

What is the difference between __getattr__ and __getattribute__? How would you implement a lazy-loaded attribute using descriptors vs __getattr__?

#
Reveal answer guidance

__getattribute__ is called unconditionally on every attribute access (before looking at instance dict). __getattr__ is only called when __getattribute__ raises AttributeError. This means __getattr__ is a fallback — it does not intercept existing attributes. For lazy loading, __getattr__ triggers computation only when the attribute does not already exist. Example: def __getattr__(self, name): if name == "expensive": value = compute(); setattr(self, name, value); return value. This works but fires only once because after setattr, the instance dict has the attribute and __getattr__ won't be called. However, this pattern fails with __slots__ or with properties. A descriptor-based approach: define a class implementing __get__ (and optionally __set__, __set_name__). On first access, __get__ computes, caches in the instance dict, and returns. Subsequent accesses hit the instance dict directly (because data descriptors with __set__ take priority over instance dicts, but non-data descriptors do not). For lazy caching with descriptors, use the cached_property decorator (Python 3.8+) — it computes once and replaces itself with the cached value in the instance dict. Descriptors are cleaner because they don't pollute __getattr__ logic and work with inheritance.

QUESTION 15PythonHard

What are Python descriptors and how do they implement @property, @staticmethod, and @classmethod internally? Why does obj.method.__func__ exist in Python 3 but not Python 2?

#
Reveal answer guidance

A descriptor is any object that defines __get__, __set__, or __delete__. Method objects are bound by a descriptor protocol: functions are non-data descriptors (only __get__). When you access obj.method, type(obj).__getattribute__ finds the function descriptor in the class dict and calls function.__get__(instance, owner), which returns a bound method object (wrapping self). @property is a data descriptor: it defines both __get__ and __set__ (if a setter is given). @classmethod is a descriptor that binds to the class (not instance): its __get__ returns a bound method with the class as first arg. @staticmethod is a descriptor whose __get__ returns the underlying function unchanged — no binding. In Python 3, obj.method.__func__ gives the original function because bound methods store both __func__ and __self__. In Python 2, unbound methods were separate types; obj.method.im_func was used. Python 3 unified functions and unbound methods — a function accessed via the class is just the function itself. Descriptors power all of Python's attribute access customization, from __slots__ (which are data descriptors) to typing.cached_property and SQLAlchemy's ORM column descriptors.

QUESTION 16PythonMedium

You deploy a Python service to production. After 24 hours, RSS grows to 2GB but sys.getsizeof() shows application objects only use ~200MB. What causes Python memory bloat in long-running services and how do you diagnose it?

#
Reveal answer guidance

Common causes: (1) Memory fragmentation — Python's allocator (pymalloc) does not return freed memory to the OS; it keeps arenas for future allocations. Use guppy3, pympler, or tracemalloc to get a full heap breakdown. (2) Cyclic references — objects referencing each other prevent reference counting from freeing them; the GC (gc module) eventually collects but defers to the gc.garbage list. gc.get_objects() lists all tracked objects. (3) C extension leaks — C extensions (NumPy, lxml, psycopg2) allocate via malloc outside Python's allocator. Use valgrind or mtrace. (4) __del__ cycles — objects with __del__ in reference cycles cannot be GC'd and leak permanently. (5) Cached data — functools.lru_cache, class-level caches, or module-level dicts accumulate entries. Middleware that caches per-request data (like SQLAlchemy identity maps) can bloat. Diagnosis: tracemalloc.start() at service start, then snapshot = tracemalloc.take_snapshot() periodically to find top allocation sites. Memory profiling with memory_profiler line-by-line. Fix: use gc.set_threshold(700, 10, 5) to tune GC, gc.collect() periodically, set PYTHONMALLOC=malloc for Valgrind, and instrument with Prometheus-style process_resident_memory_bytes tracking.

QUESTION 17PythonHard

Your team maintains a Python SDK. You need to rename get_data() to fetch_data() without breaking users. Walk through the deprecation strategy including warnings, type hints, linter integration, and removal timeline.

#
Reveal answer guidance

Strategy: (1) Add fetch_data() with the new implementation. (2) Change get_data() to call fetch_data() and emit a DeprecationWarning: warnings.warn("get_data() is deprecated, use fetch_data()", DeprecationWarning, stacklevel=2). (3) Use typing_extensions.deprecated (Python 3.13+) or the @deprecated decorator from warnings to mark in type stubs. (4) Update py.typed and add @deprecated annotation in the type stub file so IDEs and linters show strikethrough. (5) Add a Ruff or Flake8 plugin rule (RUF012 / flake8-deprecated) to flag usage. (6) In __init__.py, still export get_data to avoid ModuleNotFoundError but keep it as an alias. (7) Release notes: clearly mark as deprecated and mention the replacement in CHANGELOG under a Deprecations section. (8) Timeline: minor version 1.x — deprecation warning (no removal); major version 2.0 — remove get_data() completely. In the interim, add --no-warn-deprecated flag only for CI migration periods. (9) Linter integration: create a custom Ruff rule or use flake8-todos to flag any new usage. (10) For safety-critical SDKs, add sys.warnoptions check: if PYTHONWARNINGS=error is set, the deprecation becomes an error in CI so developers must fix before merging.

QUESTION 18PythonHard

Explain how Python's import system resolves module names, including sys.path, sys.meta_path, sys.path_hooks, and the difference between regular and namespace packages. Two packages named utils exist on sys.path — which one wins?

#
Reveal answer guidance

sys.meta_path is a list of finders. Standard finders: BuiltinImporter (built-in modules), FrozenImporter (frozen modules), PathFinder (filesystem). PathFinder iterates over sys.path (directories, zip files, namespace packages). For each path entry, it checks sys.path_importer_cache for a cached finder, else iterates sys.path_hooks to create a finder. Once a finder is found, it looks for the module in that path entry. Regular packages: __init__.py marks a directory as a package. Namespace packages (PEP 420): directories without __init__.py that share the same parent namespace — all matching directories are merged into one package. Two packages named utils on sys.path: the FIRST match in sys.path order wins. PathFinder returns a module spec from the first path entry where it finds a match (__init__.py for regular packages, or directory for namespace packages). If the first utils is a regular package, it shadows the second completely — the second is never even checked. If the first utils is a namespace package (no __init__.py), both directories are merged into one namespace package. Use importlib.util.resolve_name to see where a module resolves. sys.path order is typically: current directory first, then PYTHONPATH, then site-packages. This is why virtual environments work — site-packages come after the user's directory, so projects shadow installed packages.

CONTINUE PRACTICING

Try another perspective.