Prepare for cloud support interviews by practicing the structured diagnostic method and configuration-level reasoning employers test, kept transferable across AWS, Azure, and Google Cloud.
Real questions, real answers, the kind you could actually say out loud in an interview. Click any question to expand it.
Structured permission debugging, the single most common cloud support scenario.
I would reproduce it to see the exact error, then work through the permission layers in order: the identity's permissions, the resource policy on the bucket itself, and any conditions like region or encryption. Most access-denied issues come down to one of those, and the error message usually narrows it.
Concretely on AWS I would check the IAM policy attached to the role, then the bucket policy for an explicit deny or a condition, then whether the objects are encrypted with a key the identity is not allowed to use, because a KMS permission gap looks like a storage problem but is not. An explicit deny anywhere overrides an allow, so I look for that specifically. The same logic transfers to Azure and Google Cloud, the names change, the layered evaluation does not.
Assuming it is the bucket permissions and stopping, missing an explicit deny, a condition, or a separate encryption-key permission.
“The objects are encrypted with a customer-managed key. How does that change your check?”
Layered cloud connectivity diagnosis.
I isolate the layer. First, is the instance actually running and passing its health checks. Then the network path: is the firewall rule (security group) allowing SSH from my source, is routing correct, does it have a reachable address. Then the instance itself: is SSH running and is the key right.
The order saves time. If the instance is failing status checks, it is not a network problem, it is the instance, and I would look at the boot logs the platform exposes. If health checks pass but I still cannot connect, I focus on the security group and routing. I resist the urge to reboot first, because a reboot destroys the evidence of what was actually wrong.
Rebooting the instance immediately, which often 'fixes' it temporarily and destroys the evidence of the real cause.
“Status checks pass and the security group allows your IP, but you still cannot connect. What is left?”
The security foundation of every cloud role.
IAM, identity and access management, controls who can do what in the cloud account. Least privilege means giving an identity only the permissions it needs for its task and nothing more, so a mistake or a compromise has a limited blast radius.
In practice, least privilege is why you attach permissions to roles for specific jobs rather than handing out broad admin access. When I help a customer fix an access issue, the goal is the minimum permission that makes it work, not just adding a broad allow to make the error go away, which trades a small problem for a security hole.
“Fixing” an access issue by granting broad or admin permissions.
“A customer asks you to just give their app admin so the error goes away. How do you respond?”
Cloud networking fundamentals at the configuration level.
A security group acts like a firewall on the instance itself and is stateful, if you allow traffic in, the response is automatically allowed out. A network ACL operates at the subnet level and is stateless, so you have to allow both directions explicitly. Security groups only have allow rules; ACLs can have explicit deny rules too.
The stateful versus stateless distinction is the one that trips people up in real troubleshooting. If a subnet ACL allows inbound but forgot the outbound return traffic, connections fail in a confusing way even though the security group looks fine. In most designs the security group is the primary control and the ACL is a coarse backstop.
Forgetting that ACLs are stateless and need return traffic allowed, then chasing the wrong layer.
“Connectivity works one direction but not the other. Which of the two would you suspect and why?”
The complete Cloud Support Interview Prep Pack adds dozens more questions, deeper troubleshooting scenarios, follow-up questions, a quick-review cheat sheet, and a 7-day study plan.
DNS reasoning in a cloud context.
Raw connectivity works since the IP is reachable, so the problem is name resolution. In a cloud environment that usually means internal DNS: the private DNS zone, the resolver settings on the instance, or whether the service is registered under that name.
I would confirm with a lookup that the name is not resolving, then check whether the resource is configured to use the cloud's internal DNS, and whether the record actually exists in the private zone. It is easy to assume a firewall problem, but works-by-IP-fails-by-name almost always points at DNS, not connectivity.
Chasing security groups and routing when the symptom clearly points at name resolution.
“The name resolves from one subnet but not another. What would cause that?”
A foundational cloud concept with direct support relevance.
The cloud provider is responsible for security of the cloud, the physical infrastructure and the services themselves, and the customer is responsible for security in the cloud, their data, access controls, and configurations. The line depends on the service.
It matters in support because it tells you where a problem can and cannot live. If a customer's data is exposed, that is almost always a customer-side misconfiguration, like a storage bucket set to public, not a provider failure. Knowing the boundary helps me point the investigation at the right layer quickly and explain honestly to a customer what is theirs to fix.
Assuming the provider secures the customer's configurations and data, or blaming the platform for a customer-side misconfiguration.
“A customer's bucket was public and data leaked. Whose responsibility was that, and how do you explain it to them?”
Judgment and communication under pressure.
I would weigh active impact and risk. A production outage affecting live users and an active security issue are both urgent, and I would look at scope: how many users or how much data is affected, and whether the security issue is active exploitation or a lower-risk finding.
The part interviewers listen for is that I do not go silent on the other two. I would triage the most damaging one first, but set expectations with the others and pull in help rather than trying to solo everything. If the security issue is active exploitation, it can outrank even an outage, because the outage stops getting worse on its own while an active breach does not.
Picking one and ignoring the others, or ranking by which is technically most interesting rather than by impact.
“Midway through the outage, the security issue turns out to be active exploitation. What changes?”
Basic cloud service-model literacy.
IaaS gives you the raw infrastructure, like virtual machines, and you manage the OS and everything above it. PaaS gives you a platform to run code without managing the servers underneath. SaaS is finished software you just use.
What makes this useful rather than trivia is tying it back to shared responsibility, the more the provider manages, the less of the stack the customer is responsible for. With IaaS the customer patches the OS; with SaaS they mostly manage their data and users. The service model tells me how much of the stack is even in the customer's control to fix.
Memorizing the three terms without connecting them to who manages and secures what.
“For an IaaS virtual machine, who is responsible for patching the operating system?”
These questions come from the free SecureByDefault curriculum. Go learn the underlying skills, or explore the scenarios these answers are built on.