
Senior Cloud Infrastructure Engineer, AI Platform
Procore Technologies2 hours ago
Base Salary
$141k - $194k/yr
Responsibilities
- Implement and optimize scalable, cost-efficient, multi-tenant data ingestion and vectorization pipelines.
- Create and maintain dashboards, alerts, and logging to monitor system health and identify production bottlenecks.
- Tune performance across data retrieval, model inference, and agent response times to support millisecond-level latency.
- Design and implement secure multi-tenant infrastructure with strict tenant isolation, fair resource allocation, quotas, and cost allocation.
- Set up end-to-end observability covering logging, metrics, and alerting.
Requirements
- 5+ years of hands-on experience with cloud platforms, including strong Google Cloud Platform and Terraform expertise.
- Experience with vector databases, LLMs, and agentic observability tools, including Milvus and Arize.
- Demonstrated experience with performance analysis, tuning, infrastructure cost optimization, and explaining quantified trade-offs.
- Experience building or working on multi-tenant SaaS platforms.
- Experience with observability tools such as Prometheus, Grafana, Datadog, or GCP Operations suite.
- Preferred: understanding of challenges in training and serving large machine learning models.
- Preferred: experience with the Gemini API, LLM quota management, and load balancing.
- Preferred: hands-on experience with LLM serving frameworks and optimization techniques including quantization, tensor parallelism, and FlashAttention.
- Preferred: high-level knowledge of agentic systems and best practices.
Benefits
- Hybrid work arrangement with two days per week in the Austin office.
- Base pay range of $140,960.00–$193,820.00 USD annually.
- May be eligible for equity compensation and/or bonus incentive compensation.
- Immediate start opportunity.