Architecture you can rebuild from the repository.
Reviews, migrations, and landing zones on AWS, Azure, and GCP. The deliverable is not a diagram in a slide deck. It is the Terraform that produces the diagram, in your repository, with a plan you can run on the day we leave.
The cloud bill is usually the second problem.
Most estates we are called into were built quickly, correctly for the time, and then never revisited. One account holds everything. The database is a virtual machine somebody installed PostgreSQL on. There is a security group with 0.0.0.0/0 in it that predates the current team. The whole thing runs in a single availability zone, which is fine right up until it is not.
The spend is what gets noticed, because it arrives monthly with a number on it. But oversized instances are rarely the expensive part. The expensive part is the afternoon nobody can deploy because the one person who knows how the network was wired is on holiday, and the state of the environment exists only in the console.
So the first deliverable is a review that separates the two: what is costing money, and what is costing you the ability to change anything. They usually have the same fix, which is putting the environment into code so that a change is a pull request rather than an act of memory.
A shape that survives losing a zone.
Drawn at the level a real review happens: where the state lives, where the failure domains are, and which hops are allowed.
- Failure domain
- One AZ can go dark without a page
- State
- Managed, replicated, restore-tested
- Access
- No SSH path from the internet
- Cost
- Right-sized against real utilisation
The diagram is the readable version. What you actually receive is the Terraform that produces it, in your repository, with a plan you can run yourself.
What we take on
Landing zone and account structure
Separate accounts or subscriptions, with guardrails that apply from the top.
Production, staging, and shared services separated at the account boundary rather than by naming convention, with organisation-level policy, centralised logging, and a billing structure that tells you which team spent what. Retrofitting this later is possible, but it is always more work than doing it first.
Network and connectivity
Private subnets, controlled egress, and a path in that is not SSH from anywhere.
Application and data tiers in private subnets, egress through a gateway you can log, and administrative access through a session service or a bastion with recorded sessions. Hybrid connectivity, peering, and DNS resolution across accounts included where you have on-premise estate to reach.
Migration
A move with a rehearsal, a cutover window, and a rollback that has been tested.
Lift and shift where that is the honest answer, re-platforming where it pays for itself. Databases move to managed services with replication and a rehearsed switch, so the cutover is measured in minutes and the fallback is real rather than theoretical.
Reliability and disaster recovery
Multi-zone by default, and a recovery objective you have actually measured.
Autoscaling and health checks that remove a bad instance rather than paging a person, backups replicated across regions, and a restore performed against a stopwatch so your RTO is a number from a test rather than a number from a policy document.
Identity, secrets, and audit
Roles instead of long-lived keys, and an audit trail that is written elsewhere.
Workload identity in place of access keys checked into a repository, short-lived credentials for humans through SSO, secrets in a managed store with rotation, and audit logs shipped to an account the workload cannot write to. That last part is what makes the logs worth having after an incident.
Cost engineering
Right-sizing against real utilisation, then commitments once the shape is stable.
Utilisation measured before anything is resized, storage tiered, idle environments scheduled off, and only then commitment discounts applied, because buying a three year reservation for an instance type you are about to stop using is a common and expensive mistake.
Review, design, move, hand over.
Review
Read-only access, an inventory of what exists, and a written report of the risks, the single points of failure, and the spend.
Design
Target architecture agreed with your team, written as a document and then as Terraform modules in your repository.
Migrate
Built in staging, rehearsed, then cut over in an agreed window with a rollback that was tested rather than described.
Operate or leave
Runbooks, a walkthrough with your engineers, and the choice of holding it yourselves or keeping us on retainer.
Get the review before the renewal.
Two weeks of read-only access produces a written picture of the risks and the spend, and it is yours whether or not we do the work.