
Principal Cybersecurity Architect
Providence23 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Design, develop, optimize, and deploy Small Language Models for healthcare and cybersecurity applications.
- Implement parameter-efficient fine-tuning, instruction tuning, multi-task learning, DPO, RLHF, active learning, and continuous learning strategies.
- Build knowledge graphs, graph-augmented language models, entity extraction pipelines, relation classification systems, and graph reasoning workflows.
- Optimize inference and scalable serving with batching, caching, load balancing, A/B testing, and champion-challenger deployment patterns.
- Design vector database and hybrid retrieval architectures for semantic search and RAG applications.
- Develop advanced RAG systems with reranking, context compression, query decomposition, vector indexing, and domain-specific chunking.
- Build cyber-text data pipelines with de-identification, synthetic data generation, bias detection, quality controls, and representativeness metrics.
- Develop distributed multi-GPU and multi-node training workflows and automated data augmentation pipelines.
- Create evaluation frameworks for NER, relation extraction, summarization, question answering, reasoning, robustness, adversarial attacks, OOD detection, and hallucination detection.
- Build interpretability, monitoring, and observability tools for model performance, drift, embedding quality, retrieval accuracy, and production reliability.
- Deploy security-focused SLMs for threat detection, log analysis, vulnerability assessment, incident response automation, and policy compliance.
- Collaborate with clinical, security, product, MLOps, and engineering stakeholders to translate use cases into technical specifications.
- Mentor junior data scientists and contribute to documentation, training, technical talks, and research communities.
Requirements
- 7–10 years of experience in data science, machine learning research, or AI engineering, with a demonstrated record of developing and deploying language models.
- At least 3 years of hands-on experience with transformer architectures, fine-tuning methodologies, and production NLP systems.
- Experience applying machine learning in healthcare, life sciences, or highly regulated industries.
- Expertise in Python, PyTorch, Hugging Face Transformers, PEFT, TRL, BitsAndBytes, transformer architectures, quantization, distributed training, and memory-efficient optimization.
- Hands-on experience with LoRA, QLoRA, AdaLoRA, prefix tuning, prompt tuning, P-Tuning, IA3, adapters, GPTQ, AWQ, GGUF/GGML, and post-training quantization.
- Experience with graph databases, graph query languages, graph neural networks, knowledge graph construction, entity linking, and ontology alignment.
- Deep experience with vector databases, similarity search, embedding models, hybrid retrieval, RAG architectures, reranking, and RAG evaluation.
- Proficiency with cloud ML platforms, SQL, data warehouses, data lakes, experiment tracking, containers, orchestration, version control, and large-scale data processing.
- Knowledge of HL7, FHIR, SNOMED CT, ICD-10, LOINC, RxNorm, EHR systems, HIPAA, HITECH, de-identification standards, clinical workflows, and responsible AI.
- Strong analytical, communication, collaboration, project ownership, and research skills.
- Preferred: PhD in machine learning, NLP, computer science, or a related field, with publications in NeurIPS, ICML, ACL, EMNLP, or ICLR.
- Preferred: experience with instruction-tuning datasets, inference optimization, agentic AI frameworks, prompt optimization, clinical NLP benchmarks, multimodal models, federated learning, privacy-preserving ML, or open-source ML/AI projects.
Tech Stack
Apache AirflowApache SparkAzureDatabricksDockerDVCElasticsearchFastAPIHugging Face TransformersKerasKubernetesLightGBMMLflowNeo4jNLTKPandasPythonPyTorchReactscikit-learnSnowflakespaCyTensorFlowTerraformXGBoost