40% shorter pipelines
Re-engineered builds around Kubernetes-hosted agents and integrated code-quality and artifact tooling.
SENIOR DEVOPS & SERVICE RELIABILITY ENGINEER
CLOUD INFRASTRUCTURE · PLATFORM ENGINEERING · SRE
I design and build cloud infrastructure, Kubernetes platforms and delivery automation. I own the code, production rollout and reliability improvements that follow.

Senior DevOps Engineer · Oracle
AWS · OCI · KUBERNETES · AUTOMATIONEight years across SaaS, healthcare and e-commerce infrastructure.
My strongest platforms are AWS and OCI, complemented by hands-on Azure and GCP delivery. I develop infrastructure and deployment code, review production changes and lead migrations, upgrades and recovery.
At Oracle, I build reusable GitLab CI/CD workflows, Terraform and Python automation, release-testing tools and observability. I also develop custom Codex skills, MCP integrations and AI-assisted review workflows, with human approval retained for production changes.
Re-engineered builds around Kubernetes-hosted agents and integrated code-quality and artifact tooling.
Built reusable Terraform modules and Ansible workflows for provisioning and configuration.
Built MCP telemetry correlation and custom Codex workflows to assemble evidence and diagnostic context.
Production work is separated from public reference projects. Each case explains the implementation, choices and limits.
Terraform, Python and Bash automation for repeatable provisioning, Kubernetes lifecycle management and controlled infrastructure changes.
Develop Terraform, Python and Bash automation using OCI Resource Manager for multi-tenant platform provisioning and OKE lifecycle operations.
Led EC2-to-EKS migration at Mobile Programming and an e-commerce VM-to-GKE migration at Artisto.
Infrastructure guardrails include IAM, Vault, AWS Secrets Manager, Azure Key Vault, AWS Config, Security Hub and WAF.
Infrastructure engineering across provisioning, migration, access controls and operational readiness.
Delivery spans code review, build quality, environment readiness, deployment qualification and a deliberate response when validation fails.
Build shared GitLab CI/CD templates with security scans, linting and compliance checks. Review infrastructure and pipeline code before production promotion.
Build Playwright Release Qualification automation for integrations and schema changes, coordinating defect resolution within approved change windows.
Modernized Jenkins with containerized Kubernetes agents; integrated Maven, SonarQube, Nexus and JFrog Artifactory.
Release engineering across pipeline design, qualification, promotion and recovery.
Service symptoms, telemetry, recent changes and dependency behavior belong in one investigation—not separate dashboards.
Build Grafana SLO/SLI views, Streamlit uptime dashboards, OpenSearch/Kibana monitors and Slack-webhook notifications.
Standardize multi-region DR runbooks and failover tests for OCI storage and Oracle AI Autonomous Databases.
Missing telemetry is unknown—not healthy. Temporal correlation is a hypothesis—not proof of root cause.
Oracle outcomes: incident SLA compliance, detection and resolution are separate measures, not availability guarantees. The interactive SLO example uses synthetic counters.
A runnable incident workspace that separates deterministic findings, retrieved runbooks, uncertainty and optional LLM synthesis.
Evidence preflight, known-issue matching, incident triage, SLO analysis, UTC timelines and log-pattern grouping.
Ownership resolution, dependency impact, runbook Q&A, RCA hypotheses, code review, code-fix proposals, test plans and postmortem scaffolding.
Validate JSON → curated facts → bounded knowledge loading → BM25 retrieval → optional model → citation-ID validation.
Reference implementation using synthetic data. No production deployment, live-model evaluation, SSO or multi-tenant access-control claim.
Inspect exported Kubernetes workloads against explicit starter standards, then retrieve the policy context behind each finding.
Deployment, StatefulSet, DaemonSet and Pod exports. Resource identifiers stay attached to individual findings.
The reviewer lists skipped resource kinds and never mistakes unsupported objects for passed checks.
Deterministic starter checks provide resource-level evidence; optional RAG explains the relevant platform policy.
Reference implementation using static exports and synthetic fixtures; no live-cluster access or admission enforcement.
Identify destructive changes, internet exposure, protection removal and uncertainty without forwarding raw Terraform values to a model.
Recognize both Terraform replacement orders, deletion, forget actions and stateful-resource risks.
Reports contain resource addresses, action lists and curated findings—not raw before/after values, variables or outputs.
Retrieved rollout and recovery guidance helps a reviewer decide what evidence is required before promotion.
Reference implementation with synthetic fixtures; provider coverage is limited and the heuristic does not prove a change safe.
My public collection preserves upstream authorship and licenses. My contribution is practical operations documentation: review questions, diagnostic checks and failure-mode guidance.
Verify context, namespace and RBAC; inspect deterministic findings before requesting an LLM explanation. Treat no findings as a scope question, not a health guarantee.
Separate local validation from cloud-backed planning. Review endpoint access, IAM, node draining, networking, add-on compatibility and state recovery.
Distinguish manual pause, failed/errored/inconclusive analysis and workload readiness. Inspect traffic routing before promoting or aborting.
A completed backup is not restore proof. Review resource/storage coverage, consistency, encryption, dependencies and retention; verify recovery through application behavior.
Upstream software with my operations documentation additions. Upstream authorship and licenses are retained.
Infrastructure automation, release engineering and reliability across ECP, ECP-RT, OSDMC and ICON.
Cloud infrastructure and IoT
Develop Terraform, Python and Bash provisioning automation and OKE lifecycle workflows. Deploy LiveKit and investigate WebRTC, jitter and packet loss across Starlink, AT&T and Vodafone connectivity.
Production delivery and real-time services
Execute platform deployments and upgrades, build Playwright-based qualification and validate post-deployment behavior. Resolve infrastructure failures and coordinate application fixes within approved change windows.
Platform lifecycle and tenant operations
Develop infrastructure and delivery automation within the shared platform remit. Validate tenant-capacity changes, service health and identity integrations before and after production changes.
Infrastructure and deployment engineering
Build and review Terraform, Python, Bash and GitLab CI/CD workflows for provisioning and OKE lifecycle operations within the multi-region, multi-tenant platform environment.
04 / EXPERIENCE
Hands-on engineering across three employers.
Full résumé ↗Senior DevOps Engineer (SRE). Design and implement infrastructure automation, reusable delivery pipelines and production upgrades across multi-region, multi-tenant platforms.
DevOps Engineer. AWS and OCI healthcare platform delivery, EC2-to-EKS migration and standardized OKE rollout checks.
DevOps Engineer across cloud, platform and infrastructure. AWS foundation with Azure and GCP engagements; build automation, containers and e-commerce migration.
Deep production experience, with additional Azure and GCP engagements. Terraform · Ansible · CloudFormation · AWS CDK
EKS · OKE · AKS · GKE · Bare metal · Docker · Helm · Istio
GitLab CI/CD · Jenkins · GitHub Actions · Argo CD · OCI DevOps · Shepherd · Playwright · SonarQube
Prometheus · Grafana · OpenSearch/Kibana · ELK · OpenTelemetry · Datadog · New Relic · Streamlit
Python · Bash · Go · PowerShell · Custom Codex skills · MCP servers · AI code review
Oracle Database · MySQL/Aurora · SQL Server · Kafka · Vault · IAM · AWS Config · LiveKit · WebRTC
CONTACT