[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-evaluator-jobs-berlin-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":42},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will evaluate and rank model-generated outputs, perform QA evaluation and validation checks, follow annotation guidelines compliance requirements, and write rationales to support training data quality and model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We value strong analytical skills, attention to detail, and the ability to follow guidelines; preferred candidates may have prior data labeling, RLHF, or LLM evaluation exposure.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both German and English is required, as tasks involve bilingual evaluation and prompt evaluation.","What languages are required?",{"A":20,"Q":21},"You may evaluate a range of domains such as general knowledge, writing quality, reasoning, instruction-following, and content safety labeling scenarios, depending on project needs.","What domains are covered?","ai-evaluator-jobs-berlin-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (German\u002FEnglish) AI Generalist Trainers to support RLHF, large language model evaluation, and training data quality work by evaluating, ranking, and QA-checking model outputs to drive model performance improvement.","Germany-Based English & German AI Generalist Trainer (Remote, Full-Time) 2026 May",[27,30,33,36,39],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will contribute to RLHF and large language model evaluation by assessing model-generated responses, ranking alternatives, and writing clear rationales. You will follow annotation guidelines compliance requirements, perform QA evaluation and validation checks, and help improve training data quality to support model performance improvement across real-world use cases. This is a remote, full-time role designed for detail-oriented contributors who can reason about language, accuracy, safety, and helpfulness in both German and English.",{"h2":31,"desc":32},"Key Responsibilities","Perform large language model evaluation by reviewing prompts and model outputs in German and English; rank and compare multiple responses using defined rubrics and prompt evaluation criteria; write concise, evidence-based rationales that demonstrate sound reasoning; execute QA evaluation workflows, including consistency checks, error identification, and validation of edge cases; apply data labeling and content safety labeling standards to ensure training data quality; follow annotation guidelines compliance requirements and document issues, ambiguities, and improvement suggestions; collaborate asynchronously with leads to calibrate decisions and maintain high-quality evaluation throughput.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany; fluent in German and English (written and reading comprehension required); strong analytical skills with the ability to evaluate nuanced language and reasoning; exceptional attention to detail and consistency when following annotation guidelines; comfortable working with web-based tooling and structured feedback formats; able to meet quality targets through careful evaluation, ranking, QA, and validation.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in AI data labeling, prompt evaluation, RLHF, or QA evaluation; familiarity with LLM behavior, common failure modes, and large language model evaluation concepts; experience applying rubrics and writing rationales for ranking tasks; self-driven, dependable, and able to work independently in a remote environment while maintaining training data quality standards.",{"h2":40,"desc":41},"Pay","Compensation is $35–$40 USD per hour. If you are Germany-based and fluent in English and German, apply to Rexzone to help improve training data quality and support model performance improvement through RLHF and large language model evaluation.","AI Data Operations"]