[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-generalist-trainer-remote-jobs-germany-hiring-now":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as evaluating and ranking model outputs, conducting QA evaluation, validating labels, applying annotation guidelines compliance, and writing reasoning-based rationales.","What tasks will I do?",{"A":14,"Q":15},"AI experience is preferred but not required. Strong analytical skills, attention to detail, and the ability to follow evaluation rubrics are essential; training and calibration are provided.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required for bilingual evaluation, rationale writing, and consistency checks.","What languages are required?",{"A":20,"Q":21},"You may evaluate a wide range of domains (general knowledge, productivity, customer support-style queries, and safety-sensitive content), including content safety labeling and training data quality checks to support model performance improvement.","What domains are covered?","ai-generalist-trainer-remote-jobs-germany-hiring-now",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support RLHF and large language model evaluation by assessing, ranking, and validating model-generated outputs to drive training data quality and model performance improvement.","Germany-Based English & German AI Generalist Trainer (Remote, Full-Time) 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will help improve AI\u002FLLM workflows through RLHF-style evaluation, prompt evaluation, and QA evaluation. You will review model responses in both English and German, rank alternative outputs, validate correctness and safety, and write clear rationales that reinforce annotation guidelines compliance. Your work directly impacts training data quality, large language model evaluation, and measurable model performance improvement.",{"h2":31,"desc":32},"Key Responsibilities","Perform large language model evaluation by reviewing model-generated answers for accuracy, helpfulness, reasoning quality, and safety; rank and compare multiple candidate outputs and select best responses using defined rubrics; conduct QA evaluation to validate labels, rationales, and edge-case handling; write concise, evidence-based rationales in English and German to support evaluation decisions and reasoning; apply annotation guidelines compliance consistently, escalating ambiguous cases and proposing guideline improvements; validate content safety labeling decisions and ensure policy adherence across sensitive topics; monitor training data quality trends, identify systematic errors, and recommend workflow changes to support model performance improvement; collaborate asynchronously with operations and QA to resolve disagreements, calibrate scoring, and maintain high inter-annotator agreement.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany and authorized to work as an independent contractor or employee as applicable; fluent in both English and German (reading and writing) with strong grammar and clarity; strong analytical skills with the ability to evaluate reasoning, consistency, and factuality; exceptional attention to detail and ability to follow annotation guidelines compliance; comfortable working with web-based annotation tools and structured evaluation rubrics; able to manage time independently in a remote setting while meeting quality and throughput targets.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in data labeling, prompt evaluation, QA evaluation, or RLHF-related tasks; familiarity with LLM behavior, common failure modes (hallucinations, instruction-following issues), and evaluation methodologies; experience writing high-quality rationales and documenting decisions for audits; self-driven and proactive communicator who can flag risks, propose process improvements, and maintain consistency under shifting requirements.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour (based on assessment performance and task complexity).",{"h2":43,"desc":44},"How to Apply","Apply through Rexzone with your updated resume\u002FCV. If selected, you will complete a short skills assessment focused on bilingual evaluation, ranking, and rationale writing aligned to training data quality standards.","AI Data Operations"]