
SRE Platform Engineer
CACI International1 day ago
Base Salary
$115k - $252k/yr
Responsibilities
- Monitor and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms while maintaining availability, performance, service health, and SLO adherence.
- Implement observability through alerting, dashboards, health checks, synthetic monitoring, and log analysis using Azure Monitor, Application Insights, Log Analytics, or equivalent tools.
- Lead incident response, root cause analysis, defect resolution, dependency updates, integration validation, and emergency change coordination.
- Analyze application, API, AI model endpoint, data pipeline, and infrastructure metrics to remediate bottlenecks, latency, resource constraints, and efficiency issues.
- Support capacity planning, resource sizing, autoscaling, and cost optimization for compute, storage, and AI model consumption.
- Maintain backup and restore processes, disaster recovery procedures, high-availability architectures, and business continuity capabilities.
- Monitor Azure Databricks clusters, data pipelines, storage services, and analytical workloads for failures and performance degradation.
- Support pre-production validation, performance testing, runbook development, and go-live coordination for pilot applications and new capabilities.
- Provide surge support for complex technical issues, large-scale data collection analysis, analytical environment optimization, and specialized troubleshooting.
- Develop operational documentation including runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge-transfer materials.
Requirements
- Bachelor’s degree plus 15 years of experience in SRE, DevOps, platform engineering, systems administration, or a related field; approved equivalents include a master’s degree plus 12 years, 21 years with no degree, or an associate degree plus 17 years.
- Must be able to obtain an active DHS/EOD clearance as required.
- Extensive experience with SRE principles, including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering.
- Strong Azure cloud experience spanning compute, storage, networking, monitoring, and PaaS offerings, with knowledge of operational best practices.
- Strong experience with Azure Monitor, Application Insights, Grafana, Prometheus, the ELK stack, alerting, dashboards, and log aggregation.
- Demonstrated ability to troubleshoot complex issues across application, platform, and infrastructure layers.
- Preferred experience operating AI/ML platforms, Azure OpenAI, Databricks, Synapse, or high-scale cloud applications in production.
- Preferred experience with Azure Government, AWS GovCloud, compliance monitoring, security operations, federal operational requirements, or 24/7 mission-critical environments.
Benefits
- Flexible time off, healthcare, wellness, financial, retirement, family support, continuing education, learning and development, and other time-off benefits are offered.
- The position may be worked in more than one location; the stated salary range is a national average.
About CACI International
CACI International is a public government contractor that delivers enterprise IT, cybersecurity, electronic warfare, space, and intelligence solutions to U.S. defense, intelligence, and federal civilian customers. It provides systems engineering, software and data services, and mission support across C4ISR and business systems, primarily through long-term government contracts. Founded in 1962 and headquartered in Reston, Virginia, CACI trades on the NYSE under the ticker CACI.