[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-freelance-ai-generalist-trainer-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as prompt evaluation, ranking and preference judgments (RLHF-style), QA evaluation, validation of outputs against annotation guidelines, and writing rationales that explain your reasoning.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We value strong analytical skills, attention to detail, and consistent annotation guidelines compliance; training is provided for project-specific workflows.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in English and German is required, including the ability to read, write, and evaluate nuanced text in both languages.","What languages are required?",{"A":20,"Q":21},"Domains commonly include general knowledge, writing quality, instruction-following, content safety labeling, and other areas relevant to training data quality and model performance improvement in AI\u002FLLM systems.","What domains are covered?","freelance-ai-generalist-trainer-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support RLHF and large language model evaluation through structured ranking, QA evaluation, and rationale writing to drive training data quality and model performance improvement.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate and improve AI\u002FLLM workflows by assessing model-generated outputs, performing RLHF-style preference ranking, and documenting clear reasoning. Your work directly supports training data quality, annotation guidelines compliance, and model performance improvement by producing high-quality evaluations used in large language model evaluation pipelines.",{"h2":31,"desc":32},"Key Responsibilities","Perform large language model evaluation by reviewing English and German prompts and responses; rank and compare multiple model outputs using RLHF-style preference judgments; execute QA evaluation to validate consistency, factuality, tone, safety, and policy adherence; write concise rationales that explain reasoning behind rankings and decisions; apply annotation guidelines compliance across tasks and escalate ambiguities; validate edge cases, run spot checks, and correct labeling errors to protect training data quality; track recurring issues and provide feedback that supports model performance improvement.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany and authorized to work from Germany; fluent in English and German (reading and writing) with strong command of grammar and nuance; strong analytical skills with the ability to evaluate competing responses and justify choices with clear reasoning; high attention to detail and ability to follow annotation guidelines compliance consistently; comfortable working independently in a remote setting while meeting quality and throughput targets.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience with data labeling, content evaluation, QA evaluation, or annotation work; familiarity with LLM evaluation, prompt evaluation, or RLHF concepts; experience with content safety labeling and policy-based review; self-driven, reliable, and able to handle ambiguous examples by seeking clarifications and proposing guideline improvements.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour (hourly).",{"h2":43,"desc":44},"How to Apply","Apply through Rexzone with a short summary of your bilingual English\u002FGerman experience and any evaluation, QA, or annotation background. Highlight examples of structured reasoning, ranking decisions, and guideline-driven review work relevant to large language model evaluation.","AI Data Operations"]