[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-evaluator-jobs-frankfurt-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":48},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a full-time remote role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as evaluation, ranking, QA evaluation, validation, prompt evaluation, and writing rationales to support RLHF and training data quality.","What tasks will I do?",{"A":14,"Q":15},"AI experience is preferred but not required. You must be able to follow annotation guidelines compliance, apply strong reasoning, and maintain high training data quality.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required, including reading and writing in both languages.","What languages are required?",{"A":20,"Q":21},"Tasks span general knowledge and everyday user scenarios, including multilingual writing quality, factuality checks, instruction-following, content safety labeling, and other areas relevant to model performance improvement.","What domains are covered?","ai-evaluator-jobs-frankfurt-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based bilingual (English\u002FGerman) AI Generalist Trainers to support RLHF and large language model evaluation by ranking and validating model outputs, strengthening training data quality, and driving model performance improvement through annotation guidelines compliance.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42,45],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will contribute to AI\u002FLLM workflows by evaluating, ranking, and quality-checking model-generated responses for large language model evaluation and RLHF. You will apply annotation guidelines compliance to produce high-quality labels and rationales that improve training data quality and enable measurable model performance improvement. This is a full-time remote role for candidates located in Germany with strong bilingual fluency in English and German.",{"h2":31,"desc":32},"What You Will Do","You will assess and compare model outputs, select the best response, and write clear rationales grounded in policy and reasoning. You will perform QA evaluation, validate edge cases, and follow structured annotation guidelines to support data labeling, prompt evaluation, and content safety labeling. Your work directly supports training data quality and consistent large language model evaluation across multiple domains.",{"h2":34,"desc":35},"Responsibilities","Evaluate and rank model-generated outputs in English and German; perform QA evaluation and validation checks to ensure training data quality; write concise rationales explaining ranking decisions using sound reasoning; execute prompt evaluation and response comparison tasks for RLHF workflows; apply annotation guidelines compliance and escalate ambiguities or policy conflicts; label and review content for content safety labeling and policy adherence; track errors, identify patterns, and propose guideline clarifications that support model performance improvement; collaborate asynchronously with operations and QA to meet quality and throughput targets.",{"h2":37,"desc":38},"Basic Qualifications","Must be based in Germany and authorized to work as an independent remote contributor where applicable; fluent in English and German (reading, writing, and comprehension); strong analytical skills with the ability to evaluate nuanced responses and detect subtle errors; exceptional attention to detail and consistency when following annotation guidelines compliance; comfortable working with web-based labeling tools and structured rubrics; able to produce clear written rationales and documentation.",{"h2":40,"desc":41},"Preferred Qualifications","Prior experience in data labeling, QA evaluation, RLHF, or large language model evaluation; familiarity with LLM behavior, common failure modes, and prompt evaluation methods; experience with content safety labeling or policy-based moderation frameworks; self-driven, reliable, and able to manage time effectively in a remote environment; interest in continuous improvement and contributing feedback to improve training data quality and model performance improvement.",{"h2":43,"desc":44},"Compensation","USD $35–$40 per hour (hourly). Rate depends on assessment performance, domain fit, and ongoing quality metrics.",{"h2":46,"desc":47},"How to Apply","Apply through Rexzone with your English\u002FGerman background details and any relevant evaluation, annotation, or QA experience. If selected, you will complete a short qualification assessment focused on ranking, reasoning, and annotation guidelines compliance.","AI Data Operations"]