【JAPAN AI】AI Evaluation Scientist / English
株式会社ジーニー
給与:800万円〜1600万円
雇用形態:正社員
勤務地:東京都
仕事内容
[Role & Expectations]
As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent quality-evaluation infrastructure.
・Research and develop evaluation metrics — scientifically define "what constitutes quality" through LLM-as-Judge calibration, reward modeling, and benchmark design
・Design and build automated evaluation pipelines — integrate research outcomes into production CI/CD to deliver scalable quality gates
・Red teaming and safety verification — automate adversarial testing and build policy compliance verification frameworks
・Drive quality improvement through statistical experimental design — quantitatively verify the effectiveness of prompt strategies and model changes through A/B tests and significance testing
・Feed evaluation signals back to research and development teams — build a compound-interest loop for model improvement
・Ensure the quality of products used in production by ~200 companies through a "science of quality" approach
[Job Description]
As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent Evaluation Infrastructure.
・Evaluation Metric Research & Development
・Research and implement LLM-as-Judge calibration methods (rubric design, bias detection, proper scoring rules)
・Design, build, and validate evaluation benchmarks (construct validity, contamination detection)
・Research the application of reward modeling / preference learning to evaluation
・Select and design evaluation metrics (win rate, task success, factuality, harm detection)
・Design, build, and maintain evaluation sets (synthetic data + real logs)
・Automated Evaluation Pipeline Design & Development
・Design and implement scalable automated evaluation pipelines
・Integrate evaluation pipelines into CI/CD and build quality gates
・Design agent evaluation harnesses (multi-turn, tool use, long-context support)
・Ensure reproducibility and reliability of evaluation pipelines
・Safety & Quality Verification
・Research and implement automated red teaming (automated adversarial testing)
・Build safety and policy compliance verification frameworks
・Research and implement hallucination detection and calibration methods
・Design and execute prompt / tool regression tests
・Statistical Analysis & Experimental Design
・Design and analyze statistical experiments (A/B tests, significance testing)
・Visualize quality trends and automate regression detection
・Create quality reports and improvement proposals
・Feed evaluation signals back to research and development teams
[Key Results (KR/Metrics)]
・Evaluation coverage rate (test case coverage)
・Regression detection rate (pre-release quality degradation detection ≥ 95%)
・Evaluation pipeline execution time (completed within CI/CD)
・LLM-as-Judge and human evaluation agreement rate
・False positive / false negative rate
・Safety incident rate (post-release)
[Tech Stack]
・Languages : Python (evaluation pipelines & analysis) , TypeScript / React / Next.js (frontend) / NX
・Evaluation/QA : pytest, LangSmith, Weights & Biases, custom eval frameworks
・Data : BigQuery, Spark, Pandas
・Infrastructure : GCP (containers / K8s) , Docker, Terraform
・CI/CD : GitHub Actions
・Tools : Slack, Confluence, Linear, Google Workspace, GitHub, Notion
・AI Dev Support: Claude Code MAX Plan, Cursor, ChatGPT, Devin
・Work environment : Mac (Apple Silicon) , dual monitors available
応募資格
[You May Be a Good Fit If You]
・Education & Experience
‐Master's degree or higher (or equivalent practical experience) in Computer Science, Machine Learning, Statistics, Mathematics, Physics, Psychometrics, or related fields
‐3+ years of practical experience as an ML Engineer, Data Scientist, Research Engineer, or in ML/AI evaluation-related roles
・Technical Skills
‐Deep knowledge of LLM / generative AI evaluation methods (benchmark design, LLM-as-Judge, quantitative output quality measurement, hallucination detection, etc.)
‐Practical knowledge of statistics and experimental design (hypothesis testing, A/B testing, confidence intervals, effect sizes, etc.)
‐Experience building ML / evaluation pipelines in Python
‐Practical experience with machine learning frameworks (PyTorch, JAX, TensorFlow, etc.)
‐Experience designing and implementing evaluation metrics (task-specific metric design beyond precision/recall)
・Language requirement (at least one of the following):
‐Japanese: Fluent — able to discuss product development without friction
‐English: Business level
This position is a research and development role responsible for AI output Evaluation Science. Research or implementation experience in ML model evaluation / LLM evaluation is required.