10,000+ Active Jobs
|
500+ Hiring Companies
|
100% Verified Jobs
India
HiringGo Logo
Companies
Exclusive Jobs
Jobs Login
Homeโ€บCompaniesโ€บLH2 AI Labsโ€บFounding Software Engineer - AI Task Benchmark & Evaluation Infrastructure
LA

Founding Software Engineer - AI Task Benchmark & Evaluation Infrastructure

๐Ÿ“LOCATIONBangalore
๐Ÿ“ˆEXPERIENCE4 to 8 Yrs
๐Ÿ•˜TYPEFull time
๐ŸฅIndustryIT Services & Consulting
๐Ÿ—“POSTED13 Aug 2026

01Overview

About LH2 AI Labs Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models. We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more. For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained onthere is limited new signal left there. That is where we come in. Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today. About the roleYou will create high-quality, SWE-bench-style benchmark tasks from real-world software repositories, and build the pipelines and tooling that make task creation faster, more reliable, and more scalable. This is a dual-mandate role: roughly half hands-on task authoring and review, and half engineering the automation that the whole team depends on. Key responsibilitiesEvaluate real-world software repositories for reproducibility, test quality and suitability for benchmark task development.Understand unfamiliar architectures, dependencies and services, and establish reliable local or containerized execution environments.Identify self-contained bugs and feature changes from issues, pull requests, commits and repository history.Construct complete benchmark tasks from selected bugs and feature changes, including the task setup, specification, tests and containerized execution environment.Write precise, implementation-neutral task specifications that define the expected behavior without revealing the solution.Create and review tests that measure observable behavior, cover important edge cases and allow valid alternative implementations.Verify that the original revision demonstrates the expected failure, the reference solution passes and unrelated behavior remains stable.Independently review tasks created by other engineers and document clear, evidence-based release decisions.Improve authoring standards, validation automation and review practices while mentoring less-experienced engineers.Design and maintain the task-development pipeline: automated environment builds, validation harnesses (fail-before/pass-after checks, determinism and flakiness runs) and task packaging for evaluation at scale.Build tooling that reduces per-task authoring effort, including repository ingestion, candidate mining from issues and pull requests, containerized execution templates and CI integration for task validation.Automate quality gates so that validation evidence is generated, checked and archived without manual steps.Improve authoring standards, validation automation and review practices while mentoring less-experienced engineers. Required qualifications4+ years of professional software-engineering experience with strong debugging and code-review skills.Proficiency in at least one relevant ecosystem, such as Python, JavaScript/TypeScript, Java or Go.Hands-on experience with Git, Linux, Docker, CI/CD, dependency management and automated testing.Experience building internal tooling or CI/CD automation that other engineers depend on, such as test harnesses, build pipelines or containerized development environments.Ability to understand unfamiliar codebases, multi-service dependencies and complex build or test failures.Strong technical writing and sound judgment about whether requirements and tests evaluate behavior fairly. Preferred experienceExperience with agent-evaluation harnesses and benchmark tooling such as Harbor (agent eval framework), SWE-bench / SWE-gym style task pipelines, Terminal-Bench, or equivalent internal systems.Platform or infrastructure engineering: orchestration of containerized workloads, artifact registries, build caching or large-scale batch execution.Monorepos, microservices, open-source projects or containerized test infrastructure. What success looks likeReleased tasks are reproducible, deterministic and supported by complete validation evidence.Task specifications and tests remain aligned without leaking or over-constraining the implementation.Reviews identify defects early, maintain a low rework rate and improve team quality and throughput over time.Validation and environment setup are increasingly automated, measurably reducing time-per-released-task and eliminating manual verification steps.Tooling you build is adopted by the team and reduces onboarding time for new task authors. .

02What you'll need

Experience
4 to 8 Yrs
Employment Type
Full time
Programming languages
PythonJavaScriptJavaGoGitLinuxDockerAutomated TestingTypeScriptCICD

03About LH2 AI LABS

IT Services & ConsultingIndustry
Full timeEmployment Type
BangaloreLocation
Not Disclosed ยท salary hidden by employer
4 to 8 Yrs ยท Bangalore
Applications are reviewed directly by the hiring team.
Role Snapshot
Work ModeNot specified
Visa SponsorshipNot specified
RelocationNot specified
Job TypeFull time
LA
LH2 AI LABS
IT Services & Consulting
View all LH2 AI LABS jobs โ†’
Share