Responsibilities
- Lead deployment strategy and execution for enterprise AI infrastructure, including AI SOC, OpenShift AI, and AI-based cybersecurity log optimization services.
- Define and enforce security hardening standards, Zero-Trust architectures, and GDPR, HIPAA, and SOC2 compliance for AI deliveries.
- Develop standardized deployment blueprints and Infrastructure as Code templates for repeatable customer rollouts.
- Manage tenant isolation, software-defined networking, storage integration, and complex deployments for MSSPs and MCP servers.
- Provide highest-level technical support for critical deployment failures, performance bottlenecks, and production network connectivity issues.
- Mentor junior deployment engineers and provide technical guidance on AI stack deployment and security.
- Implement Cybersecurity AI Agents and MCP Security Implementation services.
- Optimize GPU allocation in OpenShift and Kubernetes using NVIDIA GPU Operator and MIG for large-scale inference workloads.
- Architect sovereign-cloud patterns that keep training data and inference logs within required geopolitical or organizational boundaries.
- Implement AI-service cost-governance guardrails, including GPU instance and Bedrock/Azure OpenAI token-usage monitoring.
- Establish CI/CD/CD pipelines for Cybersecurity AI Agents while preserving security hardening and network routing.
Requirements
- Bachelor’s or master’s degree in computer science, engineering, or a related field.
- 8–12 years of experience in cloud architecture, DevOps, or network security focused on large-scale infrastructure deployment.
- Deep hands-on experience deploying complex stacks across AWS, Azure, GCP, and strictly air-gapped on-premises environments.
- Expert knowledge of Red Hat OpenShift and Kubernetes for AI workloads, storage, and networking integrations.
- Experience with Zero-Trust Network Access, security hardening, and GDPR and SOC2 compliance implementation.
- Advanced Infrastructure as Code proficiency with Terraform, Ansible, or similar tools.
- Strong understanding of software-defined networking, complex routing, tenant isolation, and secure network architecture for MSSPs.
- Hands-on experience with Istio or Linkerd service meshes and mTLS communication between AI microservices and vector databases.
- Experience deploying and securing highly available, clustered vector databases such as Pinecone, Milvus, or Weaviate.
- Proficiency with Prometheus, Grafana, and OpenTelemetry for infrastructure, model-latency, and inference-drift observability.
- Experience auditing or implementing HIPAA and SOC2 controls in technical environments.
- Required or relevant certifications include AWS/GCP Solution Architect, HashiCorp Certified: Terraform Associate, and AWS/GCP Certified DevOps Engineer.
- Preferred certifications include Certified Kubernetes Administrator and AWS Certified AI Practitioner.
- Preferred experience includes AI SOC rollouts, MCP servers, or AI-based security agents.
Benefits
- Dynamic early-stage startup environment with strong customer and partner networks.
- Collaborative culture focused on innovation, continuous learning, and diversity and inclusion.
Tech Stack
About Gruve
Gruve was founded on the premise that new technologies in Machine Learning, Data Sciences, Artificial Intelligence, and Software Development are transforming Enterprise Services. Our goal is to harness these advancements to deliver services with superior efficiency and tangible outcomes. Our Team Our team is built with a strong background in Software and Services, united by a shared sense of Purpose: to achieve the best outcomes for our clients. We value all our stakeholders, recognizing that People are our most important assets. We adopt a Process framework that ensures the delivery of high-quality results every time. What Sets Us Apart Our differentiation is straightforward: we genuinely care, we innovate, we disrupt, and we work hard. Our Core Values: Customer Success: Putting customers first. Positive Feedback Loop: Embracing continuous improvement. Pursuit & Persevere: Staying resilient and ambitious. Integrity and Ethics: Acting with honesty and ethics. Team & Trust: Collaborating with trust and respect. Giving Back: Committing to community and responsibility. Gruve is Norwegian for "To Mine or Mining Activity"
