SRE / DEVOPS / PLATFORM ENGINEERING

Complex systems.
Clear thinking.
Lasting impact.

ANIL KUMAR TANGIRALA

I build the infrastructure, delivery systems and operational clarity that help teams make change with confidence.

BASED IN HYDERABAD, INDIASCROLL TO EXPLORE ↓
LIVE 3D / ILLUSTRATIVE SYSTEM MODEL
ENGINEER / BUILDER01—AKT
Professional portrait of Anil Kumar Tangirala
Anil Kumar
Tangirala

Senior DevOps Engineer · Oracle

SRE / DEVOPS / PLATFORM ENGINEERING
DESIGN WITH INTENTAUTOMATE WITH GUARDRAILSOPERATE WITH EVIDENCERECOVER WITH CONFIDENCE

Reliability isn’t
an afterthought.
It’s the architecture.

From the first Terraform module to the last recovery check, I connect the decisions that make a platform dependable.

I work across cloud architecture, infrastructure automation, release engineering and production reliability. My journey spans e-commerce, healthcare and multi-region SaaS.

Hands-on infrastructure and deployment code ownership. Thoughtful reviews. Human-validated production decisions.

Read my complete résumé ↗
01

Architecture that operates.

Access, blast radius, recovery and ownership belong in the design—not in the postmortem.

02

Delivery that earns trust.

Reusable pipelines, explicit qualification and deliberate promotion gates.

03

Evidence before certainty.

Separate observed facts, hypotheses and unknowns. Validate recovery through service behavior.

Depth beyond
the dashboard.

Production responsibilities and original reference projects.
Architecture, implementation and operational impact.

Four platforms.
Production responsibility.

My current Oracle work spans ECP, ECP-RT, OSDMC and ICON—connecting infrastructure code, controlled delivery and day-to-day reliability.

PERSONAL ENGINEERING SCOPE

Develop and review Terraform, Python, Bash and GitLab CI/CD code. Automate provisioning and OKE lifecycle operations with OCI Resource Manager across multi-region, multi-tenant platforms.

01 / CONNECTIVITYORACLE

ECP

Cloud platform & IoT connectivity

Infrastructure automation and operational support for a platform that connects cloud services with real-world devices.

  • Provisioning and OKE lifecycle automation using Terraform, Python and Bash.
  • Infrastructure and deployment-code development and review.
  • IoT connectivity troubleshooting, including jitter and packet-loss analysis across carrier and satellite networks.
Explore implementation, operations & ownership

Platform implementation

I develop Terraform, Python and Bash automation for provisioning and OKE lifecycle management. My work includes reviewing infrastructure and delivery code for multi-region, multi-tenant environments; application-code changes remain with the relevant developers.

Connectivity investigation

I troubleshoot ECP IoT connectivity across carrier and satellite networks and automate jitter and packet-loss analysis. The focus is on distinguishing network behavior from cloud or workload symptoms, rather than treating every failure as a deployment issue.

Real-time service support

I deploy LiveKit and troubleshoot WebRTC across ECP and ECP-RT. This complements the infrastructure work with service-level investigation.

Release qualification & lifecycle operations

Execute release qualification across cloud services, infrastructure components and edge artifacts, combining pipeline dry runs with tenant-lifecycle validation to prepare changes for controlled promotion.

Database change operations

Carry out procedure-led database migration and credential-rotation operations alongside platform image upgrades. Validate release prerequisites through health analysis, credential-expiry checks and monitoring-rule maintenance, coordinating application-level changes with the relevant owners.

OCI / OKETerraformIoT diagnostics
02 / REAL-TIME SERVICESORACLE

ECP-RT

Real-time communications & release operations

Production delivery and qualification alongside hands-on LiveKit and WebRTC deployment and troubleshooting.

  • Execute platform releases and upgrades across environments.
  • Perform post-upgrade activation and qualification checks; coordinate defect resolution.
  • Deploy LiveKit and investigate real-time communication behavior.
Explore implementation, operations & ownership

Deployment execution

I participate in platform deployments and upgrades across environments, with responsibility for release execution and qualification. Release completion includes qualification, not simply a successful pipeline status.

Qualification and coordination

Post-upgrade activation checks and qualification tasks expose integration issues that require coordination with product engineering. I investigate failures and work with developers on defect resolution in partnership with the application owners.

Media-service operations

I deploy LiveKit and troubleshoot WebRTC behavior. Jitter and packet-loss analysis connects the user-facing communication symptom to the underlying network and service evidence.

Maintenance beyond feature releases

Execute patch upgrades and multi-region production deployments, verify scheduled OCI compute maintenance, and carry releases through post-upgrade activation qualification. Connect infrastructure lifecycle work with service readiness rather than treating pipeline completion as the end of a release.

Release qualificationLiveKitWebRTC
03 / SERVICE OPERATIONSORACLE

OSDMC

Platform automation & operational monitoring

Cloud infrastructure engineering backed by recurring monitoring checks and production-service investigation.

  • Provisioning and OKE lifecycle work within the four-platform infrastructure scope.
  • Verify detection-rule status as part of operational monitoring.
  • Handle service alerts, including identity-service connectivity investigation.
Explore implementation, operations & ownership

Infrastructure contribution

Develop and review infrastructure and deployment automation using Terraform, Python, Bash and OCI Resource Manager, supporting provisioning and OKE lifecycle operations within the shared platform-engineering remit.

Monitoring and daily operations

Verify detection rules as part of recurring service operations, complementing dashboards with checks on the monitoring controls that support incident detection.

Incident investigation

Investigate identity-service connectivity alerts, correlate the available operational evidence and coordinate recovery with the relevant service owners.

Tenant capacity & change validation

Perform tenant storage-capacity changes with pre-change validation, tenant-status checks and post-change verification. Bring the same controlled execution discipline to capacity operations as to service releases.

Identity integration troubleshooting

Investigate partner-token generation failures and identity-service connectivity alerts, using operational evidence to narrow the affected integration and coordinate the next recovery action.

OCI / OKEMonitoring checksIncident investigation
04 / PLATFORM ENGINEERINGORACLE

ICON

Infrastructure code & Kubernetes lifecycle

One of the four platforms covered by my current Oracle infrastructure and deployment engineering responsibilities.

  • Develop Terraform, Python and Bash provisioning automation.
  • Support OKE lifecycle management through OCI Resource Manager.
  • Build and review infrastructure and GitLab CI/CD pipeline code for multi-region, multi-tenant platforms.
Explore implementation, operations & ownership

Provisioning automation

Contribute Terraform, Python and Bash automation to repeatable platform provisioning through OCI Resource Manager, with infrastructure changes represented and reviewed as code.

Kubernetes lifecycle

Support OKE lifecycle management within a multi-platform engineering remit, applying infrastructure and deployment expertise to the Kubernetes environment while collaborating with application owners.

Delivery-code ownership

Build and review infrastructure and GitLab CI/CD pipeline code for multi-region, multi-tenant platforms, connecting provisioning workflows with repeatable delivery practices.

Terraform / PythonOKE lifecycleGitLab CI/CD

ENGINEERING DEPTH / SHARED ORACLE RESPONSIBILITIES

What I build, review and operate.

Eight areas of hands-on work behind the four platforms. Shared responsibilities across my Oracle role, from infrastructure implementation and design review to release execution and production recovery.

01Architecture & change review

Review the changes that cross system boundaries.

I review Terraform changes, pipeline approvals, IAM boundaries and schema migrations as part of OCI DevOps and Shepherd delivery. Custom promotion gates validate architecture integrations and detect configuration drift between staging, pre-production and production.

  • Infrastructure and deployment-code review, with explicit environment context.
  • IAM and schema-change checks before production promotion.
  • Coordination with application owners where a change crosses team boundaries.

Focus: infrastructure design review, deployment engineering and integration readiness.

02Infrastructure automation

Turn provisioning and lifecycle work into code.

I develop Terraform, Python and Bash automation with OCI Resource Manager for platform provisioning and OKE lifecycle management across ECP, ECP-RT, OSDMC and ICON.

  • Maintain repeatable infrastructure and deployment workflows for multi-region, multi-tenant platforms.
  • Build and review the code that drives platform operations.
  • Use the same engineering discipline for lifecycle work as for initial provisioning.

Deliverables: provisioning automation, infrastructure changes and Kubernetes lifecycle workflows.

03CI/CD & release execution

Carry a change from review to controlled promotion.

I build reusable GitLab CI/CD templates and shared libraries with security scans, linting and compliance checks. I own OCI DevOps and Shepherd production deployment and upgrade responsibilities within my role.

  • Use promotion gates to validate integration assumptions and configuration consistency.
  • Use Prometheus and OpenTelemetry health signals to validate deployments and trigger rollback.
  • Coordinate release activities within approved change windows.

Deliverables: reusable GitLab templates, shared libraries and controlled promotion workflows.

04Release qualification

Test integrations and schema changes before calling a release complete.

I build Playwright-based Release Qualification automation and coordinate defect resolution. The work links release execution to integration behavior and schema-change validation.

  • Automate repeatable qualification checks using Playwright.
  • Investigate qualification failures and coordinate application fixes with developers.
  • Complete releases within approved change windows.

Deliverables: RQ automation, qualification findings and coordinated release completion.

05Observability & service health

Make uptime, reliability and cloud cost visible.

I build Streamlit uptime and cloud-cost dashboards, Grafana SLO/SLI monitoring, and OpenSearch/Kibana threshold monitors with Slack webhooks.

  • Use SLO and SLI views to connect service health to reliability targets.
  • Bring logs, metrics and deployment-health signals into operational investigations.
  • Make uptime and cost information accessible through focused dashboards.

Reported result: threshold monitoring and webhook alerting reduced detection time by 45%.

06AI-assisted SRE

Accelerate investigation while retaining human judgment.

I engineered an OpenSearch and Prometheus MCP correlation server with the SRE team and built custom Codex diagnostic workflows. I also built an AI review agent for infrastructure and deployment changes.

  • Correlate log and metric evidence to support incident triage.
  • Automate environment health reporting and paging; operate GenAI services.
  • Keep human approval before production promotion and validate production actions.

Reported results: at least 75% less triage time and 50% less initial response time.

07Disaster recovery

Exercise recovery procedures, not just document them.

I standardized multi-region disaster-recovery runbooks and failover tests, validating replication and recovery across OCI block storage and Oracle AI Autonomous Databases.

  • Maintain operational standard procedures and methods of procedure.
  • Validate replication and recovery through failover testing.
  • Coordinate recovery work across the relevant service and infrastructure owners.

Deliverables: DR runbooks, failover tests and operational procedures.

08On-call & cross-team execution

Own the operational thread through investigation and recovery.

I lead on-call incident, change and problem management. Telemetry correlation and cross-stack root-cause analysis connect the immediate response with longer-term operational improvement.

  • Investigate across infrastructure, deployment and service layers.
  • Coordinate application-code changes with developers while owning infrastructure and deployment work.
  • Use operational findings to improve procedures and subsequent change execution.

Reported results: MTTD down 50%, MTTR down 30%, and incidents resolved within SLA improving from approximately 82% to 99%.

Review scope & dependenciesImplement & review codeQualify & promoteObserve & validate recovery

ACROSS MY ORACLE ROLE

Design review → release qualification → production recovery.

OCI DevOps and Shepherd promotion gates, IAM and schema-change review, reusable GitLab templates, Playwright RQ automation, SLO/SLI dashboards, MCP-assisted diagnostics, disaster-recovery runbooks and on-call leadership. Production actions retain human validation.

50%reduction in MTTD
30%reduction in MTTR
≥75%reduction in triage time
~82% → 99%incidents resolved within SLA
Read the supporting Oracle experience in my résumé ↗

From complexity
to controlled change.

01 / UNDERSTAND

Map the system.

Gather context. Identify dependencies, ownership and the operational constraints that shape the design.

02 / ENGINEER

Make it repeatable.

Implement infrastructure and deployment automation. Review assumptions, permissions and failure paths.

03 / QUALIFY

Prove the change.

Validate configuration, integrations and service health. Keep promotion gates and recovery plans explicit.

04 / OPERATE

Close the loop.

Observe production, investigate with evidence, verify recovery and feed lessons into the platform.

Every chapter.
More perspective.

From build and release foundations to multi-region platform ownership.

Complete career profile ↗

Systems thinking.
Hands-on execution.

Technology is the means.
Dependable systems are the outcome.

01 / CLOUD

Build the foundation.

AWS · OCI · Azure · GCP

02 / INFRASTRUCTURE

Define it as code.

Terraform · Ansible · Python · Bash

03 / ORCHESTRATION

Make platforms work.

Kubernetes · OKE · EKS · Helm · Docker

04 / DELIVERY

Ship with confidence.

GitLab CI/CD · GitHub Actions · Jenkins · Argo CD

05 / OBSERVABILITY

Read the signals.

Prometheus · Grafana · OpenTelemetry · OpenSearch

06 / ENGINEERING PRACTICE

Keep judgment in the loop.

SLOs · Incident response · DR · Release qualification

THE NEXT CHAPTER

Let’s build something
worth relying on.

Start a conversation
ENGINEERING / DEEP DIVE