Innodata Inc. logo

Innodata Inc.

Quality Lead, Agentic AI Workflow Evaluation at Innodata Inc.

In Office - San Jose, CaliforniaFull-timeISL - Google - GGLCS - (Project Delivery)Posted 18 days ago
Apply with Pipeline

About the Role

<div class="content-intro"><p>Innodata&nbsp;(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.&nbsp;Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.&nbsp;We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.</p></div><p class="Paragraph SCXW66493682 BCX2"><strong>Scope of the Role:&nbsp;</strong></p> <p><span data-contrast="auto">We are standing up a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. Reviewers work through ambiguous, multi-step scenarios inside isolated test environments, assessing whether AI agents complete tasks safely, respect user intent and consent, and hold up under close scrutiny. The Quality Lead is the person accountable for whether that output is any good.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:276}">&nbsp;</span></p> <p><span data-contrast="auto">This is a senior individual contributor role. You will not manage the reviewers — that sits with the Engagement Manager — but you set the standard they are held to. You own the audit sample, run calibration, keep the rubric usable as real cases stress it, and train reviewers into the work. You are also the deputy: when the Engagement Manager is out, the engagement runs on you.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:276}">&nbsp;</span></p> <p><span data-contrast="auto">The quality approach here is not fully defined. We expect you to build it in partnership with the customer's quality leads, or at minimum to take what they have, run it honestly, and come back with specific recommendations for where it falls short.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:276}">&nbsp;</span></p> <p class="Paragraph SCXW66493682 BCX2"><strong>What You’ll Own:</strong></p> <ul> <li><span data-contrast="auto">Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and how quality gets measured and reported</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Build that system with the customer's quality leads where none exists, and where one does, operate it and recommend concrete improvements based on what the data shows</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Re-score a sample of reviewer output as a second pass; identify error patterns rather than isolated mistakes</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Run calibration sessions: surface disagreement, work it to resolution, and document the reasoning so the outcome holds for future cases</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Maintain rubric health — flag criteria that are ambiguous, overlapping, or silent on cases the team keeps hitting, and drive revisions through the customer</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Train and onboard new reviewers, including nesting plans, ramp criteria, and the judgment call on when someone is production-ready</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Give the Engagement Manager the evidence behind performance conversations: who is drifting, on what, and whether coaching is working</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Report quality trends to the Engagement Manager and, alongside them, to the customer</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Deputize for the Engagement Manager on delivery operations during absences</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Maintain information security, privacy, and facility access practices required by the customer's onsite environment</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> </ul> <p><strong>You’ll Thrive in This Role If You Have:</strong></p> <ul> <li><span data-contrast="auto">Bachelor's degree or equivalent practical experience</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">4+ years in quality assurance, quality management, or senior review work within annotation, evaluation, trust and safety, or a similarly judgment-intensive domain</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Direct experience owning a quality function: you designed the audit, not just executed someone else’s</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Hands-on familiarity with agentic systems: tool use, multi-step task execution, sandboxed environments, and common failure modes</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Demonstrated ability to run calibration with peers — including holding a position under disagreement and changing it when the argument is better</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Strong written communication; able to document a scoring standard clearly enough that a reviewer can apply it and an auditor can check it</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Comfortable in spreadsheets and in a dashboarding tool, with enough Python or SQL to pull and slice your own data (you will not be asked to build interfaces)</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> <li><span data-contrast="auto">Experience training or onboarding reviewers into rubric-based work</span><span data-ccp-props="{&quot;335559739&quot;:100}">&nbsp;</span></li> </ul> <p><em>The expected hourly salary range for this position is $75-85 p/hour, based on experience, skills, and qualifications.</em></p> <p>&nbsp;</p><div class="content-conclusion"><p></p> <p class="x_elementToProof"><em>Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at&nbsp;</em><a href="https://consumer.ftc.gov/articles/job-scams." target="_blank">https://consumer.ftc.gov/articles/job-scams.</a><em>&nbsp;</em></p> <p class="x_elementToProof"><em>If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at&nbsp;</em><a href="mailto:[email protected]" target="_blank">[email protected]</a><em>&nbsp;and consider reporting it to the FTC at&nbsp;</em><a href="http://reportfraud.ftc.gov/" target="_blank">ReportFraud.ftc.gov</a><em>.</em></p> <p></p></div>