01Responsibilities
Develop end-to-end training and fine-tuning of Large Language Models (LLMs), including both open-source (e.g., Qwen, LLaMA, Mistral) and closed-source (e.g., OpenAI, Gemini, Anthropic) ecosystemsDemonstrated publication record in AI domain especially relating to text extraction and summarizationArchitect and implement GraphRAG pipelines, including knowledge graph representation and retrieval for enhanced contextual groundingDesign, train, and optimize semantic and dense vector embeddings for document understanding, search, and retrievalDevelop semantic retrieval systems with advanced document segmentation and indexing strategiesBuild and scale distributed training environments using NCCL and InfiniBand for multi-GPU and multi-node trainingApply reinforcement learning techniques (e.g., RLHF, RLAIF) to align model behavior with human preferences and domain-specific goalsCollaborate with cross-functional teams to translate business needs into AI-driven solutions and deploy them in production environmentsComply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regard to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do soRequired Qualifications:
Bachelor's degree in computer science, Machine Learning, or related field5+ years of experience in applied AI/ML with statistics, with a solid track record of delivering production-grade modelsExperience with PyTorchExperience with Hybrid NLP solutions that combine symbolic and machine learning approachesDeep expertise in: NLP, Fundamental machine learning, deep learning, transformer, state space-based architecture Azure ML and/or AWSSolid in Python coding, SQL and database queries, data preparation, and analysisExploratory Data Analysis (EDA)Embedding models (e.g., BGE, E5, SimCSE)Semantic search and vector databases (e.g., FAISS, Weaviate, Milvus)Model fusion and ensemble techniques (stacking, boosting, gating)Deep knowledge and extensive experience with Machine/Deep Learning frameworks including transformer architectures, state space models, large language models, and agentic approachesKnowledge of algorithms and techniques within a computational domain with emphasis on text processing
Preferred Qualifications:
Experience with healthcare data and medical coding systems (e.g., CPT, CM, PCS)Document segmentation and preprocessing (OCR, layout .