Each project connects an engineering claim to its repository or documented validation. These are portfolio builds.
01 / Virtualization
VMware ESXi lab: real licensing limits, and monitoring proven by an injected failure
Claim. Built a nested VMware ESXi 8.0U3e environment from scratch, then found — through four separate operations that all returned the identical API error — that the free license blocks VM cloning, OVF deployment, and even reconfiguring an existing VM through the API; Terraform's vsphere provider crashed outright for a related reason. Automated configuration and monitoring with Ansible instead, and didn't stop at "deployed": stopped a monitored service on purpose and measured how long detection actually took.
Engineering focus. Diagnosing a Hyper-V/VT-x conflict that blocked nested virtualization before the VM would even power on, an Ansible playbook that hardens SSH and deploys a Python health-check tool plus a Prometheus/Grafana stack with its datasource and dashboard provisioned as code, and a firewall rule that had been silently blocking all of it until found and fixed.
VMware ESXigovcAnsiblePrometheusGrafana
Measured result- Monitoring detection drill: target 6 minutes (committed before the test), measured 4 minutes 23 seconds. Met.
02 / Disaster recovery
Azure disaster recovery drill: VM and SQL failover
Claim. Ran two real Azure failover drills between West US and Central US and measured them: an Azure Site Recovery VM failover and an Azure SQL failover group under a live write load. RTO and RPO targets were committed to the repository before each run, and the results are published as measured, including the one target I missed.
Engineering focus. Terraform-built infrastructure, Site Recovery replication and failover, SQL failover groups, timing recovery from the recovered VM's own log, and a cost guardrail set before anything billable existed. Both labs were torn down and checked against Azure afterward.
Azure Site RecoveryAzure SQLTerraformRTO / RPO
Measured results- VM failover: RTO target 15 min, measured 2 min 34 s (met). RPO target 5 min, measured 5 min 4 s (missed by 4 seconds).
- SQL failover group: RTO target 60 s, measured 12.9 s (met). Forced failover lost about 2.3 s of writes; planned failover lost none.
03 / Disaster recovery
EKS GitOps service with a measured database recovery drill
Claim. Deployed a service to EKS on Fargate with Argo CD, backed by Aurora, then ran a real recovery drill: the database was killed mid-traffic and restored while recovery time was measured against a target set beforehand. Target RTO was 2 minutes; measured recovery was 39 seconds.
Engineering focus. Terraform and Terragrunt for the VPC and cluster, GitOps delivery with Argo CD, Aurora, and diagnosing real failures along the way, including CoreDNS pods stuck Pending on Fargate and an AWS VPC quota limit that moved the deployment to another region. Everything was torn down and verified at zero cost afterward.
EKS FargateArgo CDAuroraTerraform
Inspect the sourceThe repository records what was not built as well: Datadog was deliberately not deployed because no account existed, and the write-up says so.
04 / AWS networking
Hybrid Cloud Network Architecture
Claim. Built a Terraform-based network design using Transit Gateway and site-to-site VPN components, then tested rebuildability through destroy and reapply.
Engineering focus. Routing, IPsec connectivity, reusable infrastructure code, and fixing a GitHub Actions OIDC trust-policy error during deployment.
AWS Transit GatewayVPNTerraformOIDC
05 / Cloud foundation
NovaRetail AWS Landing Zone
Claim. Created a single-account AWS foundation with Terraform modules for networking, identity, security controls, centralized logging, and budget visibility.
Engineering focus. Modular infrastructure and guardrails. The multi-account AWS Organizations extension is a future phase, not a completed feature.
Terraform modulesIAMLoggingAWS
06 / Networking
AWS Site-to-Site VPN across two regions
Claim. Built a real AWS Site-to-Site VPN — IPsec, a Virtual Private Gateway, a Customer Gateway, static routing — connecting two simulated company networks in two different AWS regions, with no router or ISP cooperation needed since both sides are AWS. Verified with an actual ping and SSH hostname check routed through the gateway and across the tunnel, not just a status light.
Engineering focus. Terraform building both VPCs and the VPN connection end to end, Libreswan configured with the real tunnel address and pre-shared key straight from Terraform’s own outputs, and two real bugs found with tcpdump after the obvious fixes did not work: IP forwarding disabled by default, and a security group that allowed the reply direction but not the forward direction — both produced the identical symptom until isolated.
AWS Site-to-Site VPNIPsecTerraformLibreswan
Measured result- Ping through the tunnel: 4 of 4 packets, 0% loss, ~47ms RTT. SSH hostname check confirmed a genuinely different server answered, in the other region.