Cloud infrastructure · DevOps · Operations

I build Multi-Cloud infrastructure and the checks that make deployments trustworthy.

I'm George Koufie, an IT operations professional moving into Cloud and DevOps engineering. My work combines Terraform, AWS networking, delivery pipelines, and the troubleshooting habits developed while supporting enterprise users.

Portrait of George Koufie
4 featured builds with code or project evidenceEnterprise IT background in support and operationsAWS & Azure certifications
Selected work

Projects with inspectable evidence

Each project connects an engineering claim to its repository or documented validation. These are portfolio builds.

01 / Virtualization

VMware ESXi lab: real licensing limits, and monitoring proven by an injected failure

Claim. Built a nested VMware ESXi 8.0U3e environment from scratch, then found — through four separate operations that all returned the identical API error — that the free license blocks VM cloning, OVF deployment, and even reconfiguring an existing VM through the API; Terraform's vsphere provider crashed outright for a related reason. Automated configuration and monitoring with Ansible instead, and didn't stop at "deployed": stopped a monitored service on purpose and measured how long detection actually took.

Engineering focus. Diagnosing a Hyper-V/VT-x conflict that blocked nested virtualization before the VM would even power on, an Ansible playbook that hardens SSH and deploys a Python health-check tool plus a Prometheus/Grafana stack with its datasource and dashboard provisioned as code, and a firewall rule that had been silently blocking all of it until found and fixed.

VMware ESXigovcAnsiblePrometheusGrafana
Measured result
  • Monitoring detection drill: target 6 minutes (committed before the test), measured 4 minutes 23 seconds. Met.
02 / Disaster recovery

Azure disaster recovery drill: VM and SQL failover

Claim. Ran two real Azure failover drills between West US and Central US and measured them: an Azure Site Recovery VM failover and an Azure SQL failover group under a live write load. RTO and RPO targets were committed to the repository before each run, and the results are published as measured, including the one target I missed.

Engineering focus. Terraform-built infrastructure, Site Recovery replication and failover, SQL failover groups, timing recovery from the recovered VM's own log, and a cost guardrail set before anything billable existed. Both labs were torn down and checked against Azure afterward.

Azure Site RecoveryAzure SQLTerraformRTO / RPO
Measured results
  • VM failover: RTO target 15 min, measured 2 min 34 s (met). RPO target 5 min, measured 5 min 4 s (missed by 4 seconds).
  • SQL failover group: RTO target 60 s, measured 12.9 s (met). Forced failover lost about 2.3 s of writes; planned failover lost none.
03 / Disaster recovery

EKS GitOps service with a measured database recovery drill

Claim. Deployed a service to EKS on Fargate with Argo CD, backed by Aurora, then ran a real recovery drill: the database was killed mid-traffic and restored while recovery time was measured against a target set beforehand. Target RTO was 2 minutes; measured recovery was 39 seconds.

Engineering focus. Terraform and Terragrunt for the VPC and cluster, GitOps delivery with Argo CD, Aurora, and diagnosing real failures along the way, including CoreDNS pods stuck Pending on Fargate and an AWS VPC quota limit that moved the deployment to another region. Everything was torn down and verified at zero cost afterward.

EKS FargateArgo CDAuroraTerraform
Inspect the source

The repository records what was not built as well: Datadog was deliberately not deployed because no account existed, and the write-up says so.

04 / AWS networking

Hybrid Cloud Network Architecture

Claim. Built a Terraform-based network design using Transit Gateway and site-to-site VPN components, then tested rebuildability through destroy and reapply.

Engineering focus. Routing, IPsec connectivity, reusable infrastructure code, and fixing a GitHub Actions OIDC trust-policy error during deployment.

AWS Transit GatewayVPNTerraformOIDC
Inspect the source
05 / Cloud foundation

NovaRetail AWS Landing Zone

Claim. Created a single-account AWS foundation with Terraform modules for networking, identity, security controls, centralized logging, and budget visibility.

Engineering focus. Modular infrastructure and guardrails. The multi-account AWS Organizations extension is a future phase, not a completed feature.

Terraform modulesIAMLoggingAWS
Inspect the source
06 / Networking

AWS Site-to-Site VPN across two regions

Claim. Built a real AWS Site-to-Site VPN — IPsec, a Virtual Private Gateway, a Customer Gateway, static routing — connecting two simulated company networks in two different AWS regions, with no router or ISP cooperation needed since both sides are AWS. Verified with an actual ping and SSH hostname check routed through the gateway and across the tunnel, not just a status light.

Engineering focus. Terraform building both VPCs and the VPN connection end to end, Libreswan configured with the real tunnel address and pre-shared key straight from Terraform’s own outputs, and two real bugs found with tcpdump after the obvious fixes did not work: IP forwarding disabled by default, and a security group that allowed the reply direction but not the forward direction — both produced the identical symptom until isolated.

AWS Site-to-Site VPNIPsecTerraformLibreswan
Measured result
  • Ping through the tunnel: 4 of 4 packets, 0% loss, ~47ms RTT. SSH hostname check confirmed a genuinely different server answered, in the other region.

Other work: Serverless Link Shortener API ↗.

Writing

How the pieces connect

The projects above are the evidence. This is the reflection on building them in sequence — why disaster recovery on AWS led to Azure, then to VMware, then to a real hybrid network link.

LinkedIn · October 2026

From Disaster Recovery to Hybrid Cloud: My Hands-On Journey in Cloud Infrastructure Engineering

Four projects, each one leading into the next — AWS disaster recovery, Azure Site Recovery and SQL failover, a VMware ESXi lab, and an AWS Site-to-Site VPN across two regions. The piece is less about the resume line each project adds and more about what connected them: the same RTO/RPO discipline, carried from cloud to cloud to on-prem.

LinkedIn · August 2026

Building the AWS Backend for a Volunteer-Matching Platform: Serverless, Terraform, and a Bug That Taught Me Something About Lambda Layers

A serverless backend for a volunteer-matching platform — Lambda, API Gateway, DynamoDB, atomic transactions to stop two volunteers claiming the same slot at once. The real story is the bug: a shared Python module Lambda refused to import, traced back to the layer's ZIP not following the required python/ folder structure. Terraform said the deploy succeeded; the endpoints said otherwise.

Medium

How I Reduced Deployment Errors with a GitHub Actions CI/CD Pipeline on AWS EC2

Manual deployment is where human error actually lives — the wrong version tagged, a validation step skipped under time pressure. The pipeline: GitHub Actions runs tests and a Trivy security scan, builds the Docker image, deploys it to an EC2 server Terraform provisioned, then checks a real health endpoint before calling it done. Same release steps, every time, with nothing left to remember.

GitHub ActionsTerraformDockerTrivy
Capabilities

Tools I use, connected to work I can show

Hands-on focus

  • AWS infrastructure: VPC networking, IAM, EC2, serverless services, and logging.
  • Infrastructure as code: Terraform modules and automated validation.
  • Delivery: GitHub Actions, Docker, Kubernetes, and deployment checks.
  • Operations: incident triage, identity management, documentation, and troubleshooting.

Selected credentials

  • AWS Certified Solutions Architect – Associate
  • Microsoft Certified: Azure Administrator Associate
  • CompTIA Security+ and Network+
  • B.S. Cloud Computing · Western Governors University

A credential verifies study and testing; project links show implementation. The résumé has the full certification list and employment details.

Verified badges

Issued by AWS, Microsoft, and CompTIA (via Credly).

AWS Certified Solutions Architect – Associate certificate issued to George Koufie, July 12, 2025, expires July 12, 2028. Verify at aws.amazon.com/verification.Microsoft Certified: Azure Administrator Associate certificate issued to George Koufie, earned July 4, 2025, expires July 4, 2027.
Get in touch

Let's talk cloud infrastructure and operations.

I'm interested in Cloud Engineer, Cloud Operations, DevOps, and Platform opportunities where I can contribute strong support fundamentals and growing infrastructure engineering skills.