
Software Engineer (AI Infrastructure)
BigBear.ai2 years ago
Columbia, MD, USAMid Level / Senior
Responsibilities
- Design, implement, and optimize infrastructure for AI model inference at scale.
- Develop and maintain production AI services and applications, including retrieval-augmented generation and autonomous agents.
- Define solutions for ambiguous or underspecified systems and requirements.
- Drive adoption of new technologies and practices across engineering teams.
- Implement monitoring, logging, and observability for AI services.
- Automate infrastructure provisioning and configuration using infrastructure-as-code principles.
- Ensure high availability, reliability, and performance of AI platform components.
- Contribute to security best practices for AI systems and data.
- Provide technical guidance and informal mentorship to junior engineers.
Requirements
- At least 8 years of relevant experience, or a bachelor's degree in a technical discipline plus at least 4 years of experience.
- TS/SCI clearance with polygraph.
- Experience building and maintaining production systems at scale.
- Experience with high-volume web application architecture and performance optimization.
- Strong systems integration experience across diverse technologies and platforms.
- Hands-on cloud engineering experience with AWS.
- Proficiency with Kubernetes administration and deployment patterns.
- Strong Python programming skills.
- Experience implementing observability solutions using APM, OpenTelemetry, Grafana, and Prometheus.
- Familiarity with CI/CD pipelines and DevOps practices.
- Strong change management, organizational influence, communication, and collaboration skills.
- Ability to work effectively in ambiguous environments and create structure where needed.
- Preferred: experience with vLLM, LiteLLM, LangChain, vector databases, embedding systems, high-performance computing, or distributed systems.