[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-generalist-trainer-hybrid-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. The role is remote, and you must be based in Germany to be eligible.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation, rank model-generated outputs, write rationales, run QA evaluation, validate annotations, and support training data quality through annotation guidelines compliance and content safety labeling.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not strictly required. Strong analytical skills, attention to detail, and the ability to follow guidelines consistently are essential; Rexzone provides onboarding and calibration.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both German and English is required for bilingual prompt evaluation and response assessment.","What languages are required?",{"A":20,"Q":21},"You will evaluate general-purpose prompts across domains such as everyday assistance, reasoning, writing quality, factuality, and safety, with a focus on RLHF signals, training data quality, and model performance improvement.","What domains are covered?","ai-generalist-trainer-hybrid-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based AI Generalist Trainers to support large language model evaluation across English and German. You will run RLHF-style evaluation, ranking, and QA evaluation of model outputs, follow annotation guidelines compliance, and write clear rationales that improve training data quality and drive model performance improvement in real-world AI\u002FLLM workflows.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate model-generated responses, compare alternatives, and document reasoning to support large language model evaluation. Your work directly impacts training data quality and model performance improvement through consistent prompt evaluation, data labeling, and structured feedback aligned to annotation guidelines compliance.",{"h2":31,"desc":32},"Responsibilities","• Perform large language model evaluation by assessing accuracy, helpfulness, safety, and policy alignment across EN\u002FDE prompts and responses.\n• Rank multiple model outputs and provide clear, evidence-based rationales to support RLHF and preference data creation.\n• Execute QA evaluation on labeled datasets, identify guideline deviations, and validate annotations for consistency and completeness.\n• Apply annotation guidelines compliance to data labeling tasks, including content safety labeling and sensitive-topic handling.\n• Conduct reasoning-focused reviews: verify claims, check logical consistency, and flag hallucinations or unsupported statements.\n• Validate edge cases and ambiguous prompts, propose guideline clarifications, and document recurring failure patterns.\n• Collaborate asynchronously with leads to resolve disagreements, calibrate scoring, and improve training data quality.\n• Track quality metrics, follow escalation workflows, and contribute insights for model performance improvement.",{"h2":34,"desc":35},"Basic Qualifications","• Must be based in Germany and authorized to work where applicable.\n• Fluent in German and English (reading, writing, and comprehension) for bilingual evaluation.\n• Strong analytical skills with the ability to compare responses, detect subtle errors, and justify rankings.\n• High attention to detail and consistent annotation guidelines compliance.\n• Comfortable working independently in a remote environment and meeting productivity\u002Fquality targets.\n• Able to write concise, structured rationales that reflect sound reasoning and validation.",{"h2":37,"desc":38},"Preferred Qualifications","• Prior experience in AI\u002FML data labeling, RLHF, or large language model evaluation.\n• Familiarity with LLM behavior, prompt evaluation, and common failure modes (hallucination, safety issues, bias).\n• Experience performing QA evaluation and training data quality checks.\n• Self-driven, organized, and comfortable with ambiguous problems and iterative guideline updates.",{"h2":40,"desc":41},"Compensation and Work Setup","This is a full-time, remote role for candidates based in Germany. Compensation is USD $35–$40 per hour, depending on skills alignment and evaluation performance. You will receive project onboarding, evaluation rubrics, and annotation guidelines to ensure consistent training data quality.",{"h2":43,"desc":44},"How to Apply","Apply through Rexzone with your resume\u002FCV and a short note describing your bilingual English\u002FGerman experience and any relevant evaluation, QA, or annotation background. If selected, you will complete a brief calibration assessment focused on ranking, reasoning, and annotation guidelines compliance.","AI Data Operations"]