[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-german-speaking-ai-evaluator-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":48},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role for candidates based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation by assessing and ranking model outputs, completing prompt evaluation and QA evaluation, writing reasoning-based rationales, validating outputs against policies, and supporting training data quality through annotation guidelines compliance.","What tasks will I do?",{"A":14,"Q":15},"AI\u002Fannotation experience is preferred but not required. You must be able to follow guidelines precisely, apply strong analytical skills, and produce consistent evaluations that support model performance improvement.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in German and English is required, including strong reading and writing skills in both languages.","What languages are required?",{"A":20,"Q":21},"Tasks may span general knowledge, writing quality, reasoning, summarization, customer-style queries, and content safety labeling scenarios, depending on project needs.","What domains are covered?","german-speaking-ai-evaluator-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support AI\u002FLLM workflows through RLHF-style evaluations, large language model evaluation, and training data quality initiatives that drive model performance improvement.","Germany-Based English & German AI Generalist Trainer (Remote, Full-Time) 2026 May",[27,30,33,36,39,42,45],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate, rank, and QA model-generated responses to improve large language model evaluation outcomes and overall model performance improvement. You will apply annotation guidelines compliance to produce reliable training data quality signals, write clear rationales for decisions, and validate edge cases across varied domains. This is a remote, full-time role focused on RLHF-aligned feedback, prompt evaluation, and consistent quality standards.",{"h2":31,"desc":32},"Key Responsibilities","Perform large language model evaluation by reviewing and ranking model outputs against task instructions and policy; conduct prompt evaluation and QA evaluation to verify accuracy, relevance, and helpfulness; write concise, evidence-based reasoning and rationales to support rankings and corrections; validate outputs for safety, bias, and policy adherence using content safety labeling when applicable; apply annotation guidelines compliance to ensure consistent training data quality; identify error patterns and escalate ambiguous cases with documented examples; run consistency checks and spot-audits to improve training data quality and reduce label noise; collaborate asynchronously with Rexzone leads to refine annotation guidelines and improve model performance improvement.",{"h2":34,"desc":35},"Basic Qualifications","Based in Germany and authorized to work as an independent remote contributor; fluent in German and English (professional reading and writing in both); strong analytical skills with the ability to compare alternatives and justify rankings; exceptional attention to detail and consistency under guidelines; comfortable working with web tools, spreadsheets, and structured evaluation forms; ability to follow annotation guidelines compliance and meet productivity and quality targets.",{"h2":37,"desc":38},"Preferred Qualifications","Experience with data labeling, QA evaluation, or content review in AI\u002FML pipelines; familiarity with RLHF concepts, LLM evaluation, and prompt evaluation; experience writing structured rationales and performing reasoning-based validations; prior work with content safety labeling, policy interpretation, or sensitive-content review; self-driven, reliable, and able to work independently with minimal supervision in a remote setting.",{"h2":40,"desc":41},"How You Will Be Evaluated","Quality: alignment to annotation guidelines compliance, accuracy of rankings, clarity of reasoning, and consistency across tasks; Coverage: ability to evaluate varied prompts and domains; Reliability: meeting deadlines and maintaining training data quality; Impact: actionable feedback that supports model performance improvement and stronger large language model evaluation results.",{"h2":43,"desc":44},"Compensation","USD $35–$40 per hour, depending on assessment performance and task complexity. Full-time remote engagement with ongoing work based on training data quality needs and project demand.",{"h2":46,"desc":47},"Apply","If you are based in Germany and fluent in English and German, apply to Rexzone to help improve AI systems through RLHF-aligned evaluation, ranking, QA, and high-quality rationales that strengthen training data quality and model performance improvement.","AI Data Operations"]