01Key Responsibilities
Agent Orchestration: Design and implement complex agentic workflows using LangChain and LangGraph. Manage state, cyclic graphs, and decision-making loops for autonomous agents. LLM Evaluation (LLM-as-a-Judge): Build robust evaluation pipelines. specific code/prompts where a stronger model (e.g., GPT-4) acts as a judge to verify the accuracy, safety, and adherence of the application's responses against strict requirements. Model Optimization: Execute Prompt Tuning and advanced prompt engineering strategies to align model outputs with user needs. Manage invocation calls to various foundation models (OpenAI, Anthropic, Open Source). Data Integrity: Analyze input/output data streams to ensure strict adherence to project requirements and schema validation. Ownership: operate as a core independent contributor. Required Technical Skills: Deep expertise in Python and the modern AI stack. Orchestration: Proven experience with LangChain; specific experience with LangGraph for multi-agent systems is highly preferred. Evaluation: Experience implementing "LLM-as-a-Judge" patterns, or similar evaluation frameworks. Fine-Tuning: Understanding of PEFT (Parameter-Efficient Fine-Tuning) and Prompt Tuning techniques. .