
MLOps / LLMOps Engineer (Mid-Level)
Irth Solutions24 hours ago
Remote, IndiaMid Level / Senior
Responsibilities
- Operationalize training, evaluation, packaging, deployment, monitoring, and serving for ML and LLM workloads on Databricks.
- Build reusable ML/LLM jobs, workflows, deployment templates, cluster policies, and governed lakehouse patterns using Delta Lake and Unity Catalog.
- Productionize LLM and RAG pipelines, vector search, model-serving endpoints, and batch, streaming, and online inference workflows.
- Implement data contracts, quality gates, PII detection and masking, data-residency controls, lineage, access controls, and policy-as-code.
- Automate DEV-to-QA-to-PROD promotion with Databricks Asset Bundles, GitHub Actions, testing, deployment gates, and secure secrets management.
- Develop monitoring, alerting, dashboards, runbooks, on-call procedures, disaster-recovery processes, and operational SLOs for ML/LLM services.
- Monitor and optimize model quality, data and concept drift, latency, token usage, API usage, inference costs, infrastructure costs, and workload efficiency.
- Implement resource tagging, showback and chargeback reporting, budget controls, and cost anomaly detection for AI workloads.
Requirements
- Requires 3–5 years of experience in MLOps, LLMOps, ML Engineering, Data Engineering, or platform-focused ML engineering.
- Requires hands-on Databricks experience with Databricks Jobs and Workflows, Delta Lake, Unity Catalog, and Databricks SQL Warehouses.
- Requires CI/CD experience for data and ML workloads using GitHub Actions, Databricks Asset Bundles, parameterized deployments, environment promotion, and secure secrets management.
- Requires experience with data contracts, schema governance, automated data and feature validation, and Great Expectations-style or equivalent rule-based quality frameworks.
- Requires experience building observable production pipelines with metrics, dashboards, alerting, monitoring, SLOs, MTTD, and MTTR practices.
- Requires practical experience with RBAC/ABAC, Unity Catalog security, PII detection and obfuscation, private networking, data-access controls, and policy-as-code for data residency.
- Requires strong proficiency in Python and SQL, distributed computing and orchestration in Databricks/Spark environments, and production troubleshooting and incident resolution.
- Preferred qualifications include LLM/GenAI workflows, prompt engineering, RAG, LLM evaluation, AI safety, retrieval and response-quality evaluation, and latency and cost optimization.
- Preferred qualifications include geospatial analytics with PostGIS, spatial joins, spatial indexing and tiling, coordinate systems, projections, and GIS-based feature engineering.
- Preferred qualifications include Power BI integration with Databricks SQL Warehouses and semantic layers, FinOps, Databricks disaster recovery, Azure and AWS, and cloud-native security patterns.
- Nice-to-have experience includes infrastructure-risk or asset-integrity models, data and concept drift monitoring, SLOs, incident management, and regulated infrastructure or utility data.
Benefits
- Competitive compensation based on experience and qualifications.
- Medical, dental, and vision insurance.
- 401(k) plan with company match.
- Generous paid time off and company-paid holidays.
- Remote work-from-home options are available depending on role and business needs.
- Additional on-call compensation for eligible shifts.
Tech Stack
Categories
About Irth Solutions
Irth Solutions builds cloud-based software used by utilities and other critical-infrastructure operators to manage 811 one-call tickets, land rights, risk, and mobile field work, with reporting and analytics. Its products are sold as SaaS and include AI-driven risk modeling for pipeline asset integrity. Founded in 1985 and headquartered in Columbus, Ohio, the company is privately held and owned by Gauge Capital.