22 hours ago
Doha, QatarSenior
Responsibilities
- Review delivery handoffs and approve baseline code, prompts, architecture, and documentation against maintainability standards.
- Own tiered SLA targets, incident governance, root-cause analysis, and P1/P2 mitigation across support models.
- Monitor model performance, latency, data drift, prompt configurations, and regression behavior when LLM endpoints change.
- Classify client requests as routine maintenance or system evolution and benchmark new AI models.
- Build self-healing data pipelines, automated RAG indexing synchronization, and telemetry tooling.
- Advise delivery teams on maintainable architecture and communicate technical AI impacts to government and enterprise stakeholders.
Requirements
- Require 5+ years of experience in software engineering, MLOps, SRE, or forward-deployed engineering in data-heavy or production AI environments.
- Require advanced proficiency in Python, SQL, REST/gRPC APIs, and cloud architecture using AWS, Azure, or GCP.
- Seek hands-on experience with MLOps tooling, vector databases, and LLM orchestration frameworks such as LangChain or LlamaIndex.
- Require practical knowledge of prompt version control, model benchmarking, evaluation datasets, RAG pipelines, and data-drift detection.
- Expect a systematic approach to automated fixes and a strong understanding of CI/CD for machine-learning pipelines.
- Require strong technical communication, client-expectation management, operational boundary control, and roadmap advisory skills.
Benefits
- For successful applicants based in Qatar, Scale will work with candidates and authorities to support the required visa and employment-permission application process.
- The company provides reasonable accommodations for applicants with disabilities and promotes an inclusive, equal-opportunity workplace.
About Scale AI
Scale AI builds data annotation services and AI development tools for enterprises and government agencies, sold as a platform and managed services. Its products include the Scale Generative AI Platform for building and evaluating agents and the Data Engine for collecting, curating, and labeling training data, including RLHF and model evaluation. Founded in 2016 and headquartered in San Francisco, the company is privately held and works across domains from computer vision to LLM applications.
