Actively Hiring

Senior AI Platform Engineer – MLOps & LLM Serving - Chennai (Hybrid)

Location

Chennai, Tamil Nadu / Remote (Global), Global

Work Mode

Hybrid

Experience

6-12 Yrs

Salary Range

INR 23-25 LPA

Job Type

Permanent

Openings

1 position

Job Description

Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.

Role: Senior AI Platform Engineer - MLOps & LLM Serving - Chennai (Hybrid)

Positions: 1

Experience: 6 to 12 years

Location(s): Remote (Global), Chennai

Type: Hybrid / Permanent

Salary: Up to INR 25 LPA

Notice Period: Immediate to 30 days


Key Responsibilities:

• Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control- plane topologies

• Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management

• Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression

• GitOps CI/CD, container registry, artifact/model versioning, environment promotion;observability and audit wiring (Splunk, Prometheus/Grafana)

• Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline

• Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable


Technical Skills:

• 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred

• Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory- bandwidth trade-offs

• GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks

• Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them

• Azure or AWS GPU compute operations experience


Strongly preferred

• NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery


Screening Questions

Please note that you will be asked to answer these questions during the application process:

  • 1.How many years of experience do you have with Kubernetes in production?
  • 2.Current location
  • 3.Current CTC (Lakhs per annum)
  • 4.Expected CTC (Lakhs per annum)
  • 5.Notice period
  • 6.Are you currently serving notice period?
  • 7.Do you have hands-on experience working with NVIDIA GPUs in production?
  • 8.Which LLM serving technologies have you worked with in production?
  • 9.Have you deployed and managed LLM inference workloads on GPUs?
  • 10.Which GitOps tools have you used?
  • 11.Do you have hands-on experience with Azure or AWS GPU environments?
  • 12.Have you worked on LLM inference performance tuning (latency, throughput, GPU utilisation)? (Mention which ones)

Skills & Technologies

Required Skills
ArgoCDAWS GPU InstancesAzure GPUDockerGPU ComputingGrafanaKubernetesLLM Inference & Performance TuningLLM Model ServingLLM Serving FrameworksNVIDIA GPUOpenShiftPrometheusRegulated Environments ComplianceSplunkTensorRT-LLMTerraformTriton Inference Server
EmailVerifyUploadDetails

Apply for this Role

Submit your details to apply for this role.