Senior Gen AI Engineer- RAG & Full Stack - Chennai
Chennai, Tamil Nadu / Remote (Global), Global
Remote, Hybrid
6-10 Yrs
INR 20-25 LPA
Permanent
1 position
Job Description
Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.
Role: Senior Gen AI Engineer- RAG & Full Stack - Chennai
Positions: 1
Experience: 6 to 10 years
Location(s): Chennai (Hybrid)
Type: Hybrid, Remote / Permanent
Salary: Up to INR 25 LPA (Best as per the fitment)
Notice Period: Immediate to 15 days
Key Responsibilities:
• RAG pipeline end to end: ingestion integration, hybrid retrieval with reranking, prompt/context strategy, citation resolution, refusal behavior
• Permission-aware retrieval: source ACL mapping (SharePoint/Entra, Confluence) to fail-closed retrieval filters; zero-leakage test suite partnership with QA
• Vector database design and operations (Milvus or pgvector): schema, metadata filters,sync, performance
• Full-stack product build: React/TypeScript chat and citation experience, Python/FastAP services, REST APIs, SSO/OIDC integration, admin configuration UI
• Evaluation-driven development: retrieval precision, faithfulness, and citation-accuracy metrics as the daily working loop; A/B testing of retrieval and prompt variants
• Latency engineering to the 3–5s first-token / ~15s complete-answer targets at concurrency
Technical Skills:
• 5+ years software engineering with 2+ years building RAG/LLM applications in production with real users and real quality metrics, not notebooks
• Deep retrieval craft: chunking strategy, embeddings, hybrid search, rerankers; you canexplain why retrieval fails and how you measured the fix
• Genuine full-stack evidence: shipped React/TypeScript front ends AND Python back-end services in production; API design; OIDC/SAML integration
• Vector database production experience (Milvus, pgvector, Weaviate, or equivalent) including permission/metadata filtering
• Evaluation fluency: has built or operated a retrieval/answer quality harness with numeric thresholds
Strongly Preferred:
• Permission-aware/multi-tenant retrieval specifically; Microsoft Graph API; NIM/OpenAI-compatible serving endpoints; streaming UX; enterprise design systems; banking content domains
Screening Questions
Please note that you will be asked to answer these questions during the application process:
- 1.How many years of experience do you have in software engineering?
- 2.Current location
- 3.Current CTC (Lakhs per annum)
- 4.Expected CTC (Lakhs per annum)
- 5.Notice period
- 6.Are you currently serving notice period?
- 7.How many years of hands-on experience do you have building and deploying RAG/LLM applications in production?
- 8.Have you worked on a RAG application used by real users in production?
- 9.Have you worked with Vector Databases in production? Please mention the database(s).
- 10.Which of the following have you worked hands-on with in production? (Chunking strategies ,Embeddings ,Hybrid search ,Reranking ,Retrieval evaluation/metrics)
- 11.Do you have production experience building both React/TypeScript frontend and Python backend services?
- 12.Which Python backend framework have you used in production?
- 13.Have you implemented permission/metadata-based filtering in a RAG or vector search system?
Skills & Technologies
Apply for this Role
Submit your details to apply for this role.