Search Jobs

Search by job, company or skills

Agent Algorithm Evaluation Engineer

Agent Algorithm Evaluation Engineer

Shopee
  • Posted 5 days ago
  • Be among the first 10 applicants

Job Description

Job Description:

  • Evaluation Framework Design: Design and build multi-dimensional evaluation metric systems tailored to diverse business scenarios.
  • Red-Team Testing & Safety Control: Simulate complex, extreme real-world user scenarios to conduct stress testing and red-teaming of AI models identify and drive remediation of issues related to bias, hallucinations, policy-violating content, and values alignment.
  • In-Depth Bad Case Analysis: Conduct root-cause analysis of erroneous model outputs and collaborate with algorithm engineers (Model/SFT/RLHF) to drive prompt optimization and fine-tuning improvements.

Automated Evaluation Tooling: Leverage LLM-as-a-Judge methodologies to build automated evaluation pipelines that enhance product iteration efficiency.

  • User Experience Insights: Conduct in-depth research into user psychology within social and companion-style products, translating subjective, qualitative experience into quantifiable metrics to continuously enhance product experience.

Requirements:

  • Proven experience enhancing automated evaluation workflows, including but not limited to conventional evaluation methods, LLM-as-a-Judge, and automated weakness mining.
  • Proven experience enhancing automated data production pipelines, including but not limited to Self-Instruct, Evol-Instruct, and Synthetic Data RL.
  • Strong software engineering fundamentals, with proficiency in Python and Shell scripting and hands-on experience in Linux environments.
  • Master's degree or above in Computer Science or a related field, with at least 3 years working experience.

Nice-to-Have

  • Prior background in AI evaluation, QA automation, or NLP algorithm development at a major internet company.
  • Extensive experience as a power user or developer of mainstream AI social products.

More Info

Key Skills

RLHF

LLM-as-a-Judge

Self-Instruct

Synthetic Data RL

Evol-Instruct

Model SFT

About Company