Senior AI Platform Engineer – MLOps & LLM Serving - Chennai (Hybrid)
Chennai, Tamil Nadu / Remote (Global), Global
Hybrid
6-12 Yrs
INR 23-25 LPA
Permanent
1 position
Job Description
Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.
Role: Senior AI Platform Engineer - MLOps & LLM Serving - Chennai (Hybrid)
Positions: 1
Experience: 6 to 12 years
Location(s): Remote (Global), Chennai
Type: Hybrid / Permanent
Salary: Up to INR 25 LPA
Notice Period: Immediate to 30 days
Key Responsibilities:
• Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control- plane topologies
• Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management
• Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression
• GitOps CI/CD, container registry, artifact/model versioning, environment promotion;observability and audit wiring (Splunk, Prometheus/Grafana)
• Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline
• Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable
Technical Skills:
• 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred
• Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory- bandwidth trade-offs
• GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks
• Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them
• Azure or AWS GPU compute operations experience
Strongly preferred
• NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery
Screening Questions
Please note that you will be asked to answer these questions during the application process:
- 1.How many years of experience do you have with Kubernetes in production?
- 2.Current location
- 3.Current CTC (Lakhs per annum)
- 4.Expected CTC (Lakhs per annum)
- 5.Notice period
- 6.Are you currently serving notice period?
- 7.Do you have hands-on experience working with NVIDIA GPUs in production?
- 8.Which LLM serving technologies have you worked with in production?
- 9.Have you deployed and managed LLM inference workloads on GPUs?
- 10.Which GitOps tools have you used?
- 11.Do you have hands-on experience with Azure or AWS GPU environments?
- 12.Have you worked on LLM inference performance tuning (latency, throughput, GPU utilisation)? (Mention which ones)
Skills & Technologies
Apply for this Role
Submit your details to apply for this role.