[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-evaluator-jobs-cologne-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a full-time remote role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation work including evaluation and ranking of model outputs, prompt evaluation, QA evaluation, validation against rubrics, and writing clear reasoning\u002Frationales while following annotation guidelines compliance.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We value strong analytical skills, attention to detail, and the ability to follow guidelines to maintain training data quality and support model performance improvement.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required because tasks involve bilingual evaluation and labeling.","What languages are required?",{"A":20,"Q":21},"Domains vary and can include general knowledge, customer-support style conversations, reasoning tasks, content safety labeling, and other real-world prompts used for RLHF and training data quality improvements.","What domains are covered?","ai-evaluator-jobs-cologne-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support large language model evaluation through RLHF-style ranking, prompt evaluation, and QA evaluation. You will label and review model outputs, write clear rationales, and follow annotation guidelines compliance to strengthen training data quality and drive model performance improvement in real-world AI\u002FLLM workflows.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate and improve AI systems by assessing, ranking, and validating model-generated outputs. Your work supports RLHF, large language model evaluation, and training data quality initiatives by applying detailed annotation guidelines, content safety labeling policies, and quality checks that directly contribute to model performance improvement.",{"h2":31,"desc":32},"Responsibilities","Evaluate and rank model responses against defined criteria; perform QA evaluation on labeled datasets for accuracy and consistency; write concise, evidence-based reasoning\u002Frationales for rankings and decisions; validate prompts, outputs, and edge cases for policy and instruction adherence; apply annotation guidelines compliance across English and German tasks; conduct prompt evaluation and error analysis to flag recurring model failure patterns; support training data quality reviews, spot-checks, and escalations for ambiguous cases; label and review content safety labeling categories (e.g., sensitive content, toxicity, privacy) as required; document decisions and maintain high-quality notes to support calibration and team alignment.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany and available for full-time remote work; fluent in both English and German (written and reading comprehension required); strong analytical skills with the ability to compare nuanced responses and justify rankings; high attention to detail and consistency when following guidelines; comfortable working with web-based annotation tools and structured rubrics; ability to produce clear written reasoning and handle repetitive evaluation tasks with sustained quality.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in data labeling, QA evaluation, content moderation, or annotation workflows; familiarity with LLM evaluation, RLHF concepts, or prompt evaluation practices; experience applying annotation guidelines at scale with calibration and feedback loops; self-driven, dependable, and able to manage time effectively in a remote environment; interest in improving model behavior through training data quality and systematic validation.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour (hourly).",{"h2":43,"desc":44},"How to Apply","Apply to Rexzone with a brief summary of your English\u002FGerman proficiency, Germany location, and any experience relevant to large language model evaluation, data labeling, and QA evaluation. Qualified candidates may complete a short assessment focused on ranking, reasoning, and annotation guidelines compliance.","AI Data Operations"]