
SRE Platform Engineer
CACI International1 day ago
Base Salary
$115k - $252k/yr
Responsibilities
- Monitor and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms while meeting availability, performance, and SLO objectives.
- Implement observability through alerting, dashboards, health checks, synthetic monitoring, and log analysis using Azure Monitor, Application Insights, Log Analytics, or equivalent tools.
- Lead incident response, troubleshooting, root cause analysis, defect resolution, dependency updates, integration validation, and emergency changes.
- Analyze application, API, AI model endpoint, data pipeline, and infrastructure metrics to resolve bottlenecks, latency, resource constraints, and efficiency issues.
- Support capacity planning, resource sizing, autoscaling, and cost optimization for compute, storage, and AI model consumption.
- Maintain backup and restore processes, disaster recovery procedures, high-availability architectures, and business continuity capabilities.
- Monitor Azure Databricks clusters, data pipelines, storage services, and analytical workloads for failures and performance degradation.
- Support pilot application deployments and operational readiness through pre-production validation, performance testing, runbook development, and go-live coordination.
- Provide surge support for complex technical issues, large-scale data collection analysis, analytical environment optimization, and specialized troubleshooting.
- Develop runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge-transfer materials.
Requirements
- Bachelor’s degree and 15 years of experience in site reliability engineering, DevOps, platform engineering, systems administration, or a related field; equivalencies include a master’s degree and 12 years, 21 years with no degree, or an associate degree and 17 years.
- Must be able to obtain an active DHS/EOD clearance as required.
- Extensive experience with SRE principles, including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering.
- Proven expertise with Azure cloud services, including compute, storage, networking, monitoring, and PaaS offerings.
- Strong experience with Azure Monitor, Application Insights, Grafana, Prometheus, the ELK stack, alerting, dashboards, and log aggregation.
- Demonstrated ability to troubleshoot complex issues across application, platform, and infrastructure layers.
- Experience with AI/ML platforms, Azure OpenAI, Databricks, Synapse, or high-scale cloud applications is desired.
- Experience with Azure Government or AWS GovCloud, compliance monitoring, security operations, federal environments, or 24/7 mission-critical operations is desired.
Benefits
- Flexible time off, healthcare, wellness, financial and retirement benefits, family support, continuing education, and other time-off benefits are provided.
- Access to learning and development resources and opportunities for continuous career growth.
- The position can be worked in more than one location; no travel is required.
About CACI International
CACI International is a public government contractor that delivers enterprise IT, cybersecurity, electronic warfare, space, and intelligence solutions to U.S. defense, intelligence, and federal civilian customers. It provides systems engineering, software and data services, and mission support across C4ISR and business systems, primarily through long-term government contracts. Founded in 1962 and headquartered in Reston, Virginia, CACI trades on the NYSE under the ticker CACI.