岗位描述
职位描述
- AI Quality Assessment Framework & Standards
- Establish and maintain quality evaluation criteria for LLMs, covering accuracy, safety, compliance, and instruction-following capabilities.
- Design evaluation workflows combining automation and human-in-the-loop assessment to quantify model performance regularly.
- End-to-End Quality Control & Risk Management
- Own regression testing and acceptance testing for model iterations, enforcing quality gates to prevent performance degradation before launch.
- Monitor online performance to mine for badcases, identifying model hallucinations, logic errors, or experience defects, and issue timely risk alerts.
- Issue-Driven Improvement & Closed-Loop Management
- Lead Root Cause Analysis for quality issues, precisely diagnosing whether failures stem from data, prompts, or model architecture.
- Drive algorithm, product, and operations teams to resolve quality defects, track the fix rate and conduct re-validation to ensure a closed-loop process.
- Cross-functional Collaboration
- Bridge the gap between business needs and technical implementation, ensuring the AI delivers real value to users.
任职资格
- Educational Background: Full-time bachelor's degree or higher.
- Work Experience: 3+ years of experience in internet product operations. Hands-on experience in LLM/AIGC product operations is highly preferred.
- Professional Skills:
- Mastery of AI evaluation methodologies (e.g., Golden Dataset construction, Human Evaluation SOPs).
- Sharp eye for detail ("bug hunting" mindset) with the ability to mine potential risks from massive data.
- Familiarity with Prompt Engineering to reproduce or verify issues (focus on validation rather than daily maintenance).
- Soft Skills: Strong cross-functional communication skills; ability to use data to influence and drive cross-functional partners to solve complex problems.
- Language Skills: Proficiency in written and spoken English.
- Preferred Qualifications: Project management experience is a plus.