[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-generalist-trainer-jobs-hybrid-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation, ranking and comparing model outputs, writing reasoning\u002Frationales, running QA evaluation and validation checks, and completing data labeling and content safety labeling while maintaining annotation guidelines compliance.","What tasks will I do?",{"A":14,"Q":15},"AI or annotation experience is preferred but not required. You must have strong analytical skills, attention to detail, and the ability to follow rubrics to support training data quality and model performance improvement.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required for bilingual evaluation work.","What languages are required?",{"A":20,"Q":21},"You may evaluate content across general knowledge, reasoning, writing quality, instruction-following, and safety-sensitive scenarios, supporting RLHF and broader AI\u002FLLM workflows.","What domains are covered?","ai-generalist-trainer-jobs-hybrid-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based bilingual (English\u002FGerman) AI Generalist Trainers to support RLHF and large language model evaluation by assessing, ranking, and QA-checking model outputs to drive training data quality and model performance improvement.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will contribute to AI\u002FLLM workflows by performing RLHF-style evaluation, prompt evaluation, and QA evaluation across diverse tasks. You will review model-generated responses, rank alternatives, write clear rationales, and validate outputs against annotation guidelines compliance to strengthen training data quality and enable model performance improvement. This is a remote, full-time role paid at $35–$40\u002Fhour.",{"h2":31,"desc":32},"Key Responsibilities","Perform large language model evaluation by reviewing and scoring outputs for correctness, helpfulness, safety, and policy adherence; rank multiple candidate responses and provide concise reasoning\u002Frationales aligned to rubrics; execute QA evaluation, validation, and consistency checks to improve training data quality; apply and uphold annotation guidelines compliance, including edge-case handling and escalation when needed; complete data labeling and content safety labeling for multilingual (English\u002FGerman) scenarios; run prompt evaluation to identify failure modes and propose fixes that support model performance improvement; document decisions, track recurring issues, and collaborate with team leads to refine evaluation standards.",{"h2":34,"desc":35},"Basic Qualifications","Based in Germany and able to work remotely from Germany; fluent in English and German (reading, writing, and nuanced comprehension); strong analytical skills with the ability to compare options and justify rankings; excellent attention to detail for consistent annotation guidelines compliance; comfort working with structured rubrics, examples, and QA checklists; reliable internet connection and ability to meet quality and throughput targets in a full-time schedule.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in AI data labeling, LLM evaluation, RLHF, prompt evaluation, or QA evaluation; familiarity with common LLM failure modes (hallucinations, reasoning errors, instruction-following issues); experience applying safety policies and performing content safety labeling; self-driven, organized, and able to work independently while maintaining high training data quality; comfort writing clear rationales in both English and German.",{"h2":40,"desc":41},"Skills (Role-Aligned)","RLHF, large language model evaluation, LLM evaluation, data labeling, prompt evaluation, QA evaluation, annotation guidelines, annotation guidelines compliance, content safety labeling, training data quality, ranking, reasoning, validation, rubric-based scoring, bilingual evaluation (English\u002FGerman), model performance improvement",{"h2":43,"desc":44},"How to Apply","Apply through Rexzone with a brief summary of your bilingual (English\u002FGerman) background, your Germany-based availability for a remote full-time schedule, and any relevant evaluation, QA, or annotation experience. We review applications on a rolling basis.","AI Data Operations"]