01Overview
About Client
Our client was founded by IIT Kanpur graduates (2005 batch) operates 3 verticals:
AI & ML services - building artificial-intelligence solutions for clients in the US.
Media Analytics - working with large YouTube channels and media houses.
Talent Solutions - meeting staffing and technology-hiring needs for large enterprises and mid-size companies.
The selected person is hired as a full-time employee (FTE) on our client's payroll and is deployed to work on the end client's projects.
About the Engagement
The role sits within the AI engineering team of a large global financial services enterprise, building production AI applications on top of large language models. The team's focus is retrieval-augmented systems that let business users query large internal document estates accurately and reliably, together with the services and APIs that put those capabilities into everyday use. The work is hands-on and product-facing, with early exposure to emerging agent-based architectures.
About the Role
You will design and build intelligent AI-driven applications using large language models, retrieval-augmented generation (RAG) pipelines and emerging agent-based systems. This is a hands-on engineering role - implementing scalable AI solutions end to end, contributing to architecture, and working closely with lead engineers, product managers and data scientists. It suits an engineer with genuine machine-learning and NLP foundations who has since moved into building and shipping LLM-based applications.
Key Responsibilities
AI/ML & NLP development - Design and implement ML/NLP models and pipelines; perform data preprocessing, feature engineering and model evaluation; contribute to improving model accuracy, robustness and efficiency.
LLM & generative AI - Build applications using large language models; develop and optimise prompts, embeddings and inference workflows; support fine-tuning and evaluation of LLM-based systems.
RAG pipelines - Develop and maintain retrieval-augmented generation pipelines integrating vector search with LLMs; work with embedding models and vector stores; improve retrieval quality and response relevance.
Agent-based systems - Contribute to agent-based workflows and orchestration logic; implement basic agent coordination and tool integrations; work with emerging frameworks such as LangChain agents or Google ADK.
System development & deployment - Build reusable, scalable AI services and APIs; develop solutions for real-time and batch inference; ensure code quality, performance optimisation and reliability.
Collaboration & execution - Work closely with lead engineers, product managers and data scientists; participate in design discussions and technical reviews; contribute to documentation and knowledge sharing.
Required Skills & Experience
Experience - 5 7 years of software engineering experience, including 3 years in machine learning / NLP and 1.5+ years building LLM-based applications.
RAG (Retrieval-Augmented Generation) - Hands-on experience building and improving retrieval pipelines: chunking, embeddings, retrieval quality and response relevance.
Vector databases - Practical experience with any one of Elasticsearch, OpenSearch, FAISS or a similar vector store, including embedding generation and similarity search.
Python - Strong, production-grade coding skills in Python.
LLM application development - Hands-on experience building applications with any one of GPT, Llama, Gemini, Claude or a similar model, including prompt engineering and LLM evaluation.
NLP fundamentals - Solid grounding in tokenization, embeddings, transformers and model evaluation.
APIs & deployment - Experience with REST APIs / microservices, Docker, and any one cloud platform (Azure, AWS, GCP or OCP).
Nice-to-Have / Preferred
LLM frameworks - Exposure to any one of LangChain, Google ADK or a similar framework.
Agent-based systems - Exposure to agent workflows, tool-based reasoning and orchestration patterns.
ML frameworks - Working exposure to any one of PyTorch, TensorFlow or a similar framework.
Model training & fine-tuning - Familiarity with model training and fine-tuning approaches.
MLOps & monitoring - Exposure to MLOps tooling, model monitoring and drift detection; Kubernetes an advantage.
Search & knowledge - Familiarity with hybrid search or knowledge graphs.
Domain - Financial services / BFSI exposure advantageous. .