[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-evaluator-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, but you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as evaluation and ranking of model outputs, prompt evaluation, QA evaluation, validation checks, and writing reasoning-based rationales to improve training data quality.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We provide task-specific instructions and annotation guidelines; strong analytical skills and attention to detail are essential.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required, including strong reading and writing skills for bilingual evaluation and data labeling.","What languages are required?",{"A":20,"Q":21},"Tasks may span general knowledge, business writing, customer support scenarios, content safety labeling, and other real-world prompts used for RLHF and model performance improvement.","What domains are covered?","ai-evaluator-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual English\u002FGerman AI Generalist Trainers to support RLHF and large language model evaluation by assessing, ranking, and validating model-generated outputs to strengthen training data quality and drive model performance improvement.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will work in AI\u002FLLM workflows focused on RLHF, large language model evaluation, and training data quality. You will evaluate, rank, and QA model responses, write clear rationales, and ensure annotation guidelines compliance so the team can deliver consistent training signals that support model performance improvement across multilingual use cases.",{"h2":31,"desc":32},"Responsibilities","Perform large language model evaluation by rating and ranking model-generated outputs in English and German; apply RLHF-style preference ranking and prompt evaluation to identify the best responses; conduct QA evaluation through consistency checks, edge-case reviews, and validation of annotations; write concise, evidence-based reasoning and rationales for decisions; follow annotation guidelines compliance requirements and document exceptions; execute data labeling tasks including content safety labeling when required; validate training data quality using checklists, rubrics, and targeted audits; escalate ambiguous cases and propose guideline updates to improve clarity and throughput.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany and eligible to work remotely from Germany; fluent in English and German (C1\u002FC2 or equivalent) with strong writing skills in both languages; strong analytical skills with the ability to compare outputs and justify rankings using clear reasoning; exceptional attention to detail and consistency when applying rubrics and annotation guidelines; comfortable working with structured evaluation tasks, spreadsheets\u002Ftools, and written feedback loops.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in data labeling, QA evaluation, prompt evaluation, or RLHF-related workflows; familiarity with LLM behavior, common failure modes, and evaluation criteria for helpfulness, correctness, and safety; experience with annotation guidelines creation or improvement and training data quality initiatives; self-driven, reliable, and able to manage time effectively in a remote, high-accuracy environment.",{"h2":40,"desc":41},"Compensation","Pay is $35–$40 USD per hour, based on skills and assessment performance.",{"h2":43,"desc":44},"How to Apply","Apply to Rexzone with a short summary of your English\u002FGerman language background and any experience in evaluation, QA, data labeling, or content review. If shortlisted, you will complete a structured assessment covering ranking, reasoning, and annotation guidelines compliance.","AI Data Operations"]