Actively Hiring

Senior Gen AI Engineer- RAG & Full Stack - Chennai

Location

Chennai, Tamil Nadu / Remote (Global), Global

Work Mode

Remote, Hybrid

Experience

6-10 Yrs

Salary Range

INR 20-25 LPA

Job Type

Permanent

Openings

1 position

Job Description

Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.

Role: Senior Gen AI Engineer- RAG & Full Stack - Chennai

Positions: 1

Experience: 6 to 10 years

Location(s): Chennai (Hybrid)

Type: Hybrid, Remote / Permanent

Salary: Up to INR 25 LPA (Best as per the fitment)

Notice Period: Immediate to 15 days


Key Responsibilities:

• RAG pipeline end to end: ingestion integration, hybrid retrieval with reranking, prompt/context strategy, citation resolution, refusal behavior

• Permission-aware retrieval: source ACL mapping (SharePoint/Entra, Confluence) to fail-closed retrieval filters; zero-leakage test suite partnership with QA

• Vector database design and operations (Milvus or pgvector): schema, metadata filters,sync, performance

• Full-stack product build: React/TypeScript chat and citation experience, Python/FastAP services, REST APIs, SSO/OIDC integration, admin configuration UI

• Evaluation-driven development: retrieval precision, faithfulness, and citation-accuracy metrics as the daily working loop; A/B testing of retrieval and prompt variants

• Latency engineering to the 3–5s first-token / ~15s complete-answer targets at concurrency


Technical Skills:

• 5+ years software engineering with 2+ years building RAG/LLM applications in production with real users and real quality metrics, not notebooks

• Deep retrieval craft: chunking strategy, embeddings, hybrid search, rerankers; you canexplain why retrieval fails and how you measured the fix

• Genuine full-stack evidence: shipped React/TypeScript front ends AND Python back-end services in production; API design; OIDC/SAML integration

• Vector database production experience (Milvus, pgvector, Weaviate, or equivalent) including permission/metadata filtering

• Evaluation fluency: has built or operated a retrieval/answer quality harness with numeric thresholds


Strongly Preferred:

• Permission-aware/multi-tenant retrieval specifically; Microsoft Graph API; NIM/OpenAI-compatible serving endpoints; streaming UX; enterprise design systems; banking content domains

Screening Questions

Please note that you will be asked to answer these questions during the application process:

  • 1.How many years of experience do you have in software engineering?
  • 2.Current location
  • 3.Current CTC (Lakhs per annum)
  • 4.Expected CTC (Lakhs per annum)
  • 5.Notice period
  • 6.Are you currently serving notice period?
  • 7.How many years of hands-on experience do you have building and deploying RAG/LLM applications in production?
  • 8.Have you worked on a RAG application used by real users in production?
  • 9.Have you worked with Vector Databases in production? Please mention the database(s).
  • 10.Which of the following have you worked hands-on with in production? (Chunking strategies ,Embeddings ,Hybrid search ,Reranking ,Retrieval evaluation/metrics)
  • 11.Do you have production experience building both React/TypeScript frontend and Python backend services?
  • 12.Which Python backend framework have you used in production?
  • 13.Have you implemented permission/metadata-based filtering in a RAG or vector search system?

Skills & Technologies

Required Skills
ConfluenceEmbeddingsFastAPIFull Stack DevelopmentGenAILarge Language Models (LLM)Microsoft Entra IDMicrosoft Graph APIPythonRAG (Retrieval-Augmented Generation)RAG EvaluationReactRerankingSharePointStreamingTypeScriptVector DB
EmailVerifyUploadDetails

Apply for this Role

Submit your details to apply for this role.