[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-llm-evaluator-jobs-hamburg-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a full-time remote role, but you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as evaluation and ranking of model outputs, QA evaluation of labeled data, prompt evaluation, validation of edge cases, and writing clear rationales to support training data quality and model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We value strong analytical skills, attention to detail, and the ability to follow annotation guidelines compliance; you will be assessed on practical evaluation and reasoning tasks.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required, including the ability to read, evaluate, and write rationales in both languages.","What languages are required?",{"A":20,"Q":21},"Domains vary and can include general knowledge, reasoning, writing quality, translation, instruction following, and content safety labeling, all aligned to RLHF and training data quality goals.","What domains are covered?","llm-evaluator-jobs-hamburg-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support RLHF and large language model evaluation by ranking, QA evaluation, and writing clear rationales that drive training data quality and model performance improvement.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-Based English & German AI Generalist Trainer at Rexzone, you will evaluate and improve AI\u002FLLM workflows by assessing model-generated outputs, ranking alternatives, and documenting reasoning. Your work directly supports RLHF, large language model evaluation, and training data quality through consistent application of annotation guidelines compliance, content safety labeling, and training data quality checks.",{"h2":31,"desc":32},"Key Responsibilities","Evaluate and compare model-generated answers in English and German using defined rubrics; Rank multiple outputs and provide defensible reasoning and concise rationales; Perform QA evaluation on labeled datasets to ensure training data quality and annotation guidelines compliance; Validate edge cases, identify failure modes, and flag policy or content safety labeling issues; Apply prompt evaluation and LLM evaluation methods to measure helpfulness, correctness, tone, and safety; Review and reconcile disagreements to improve consistency and inter-annotator alignment; Document decisions, follow workflow instructions, and escalate unclear guidelines for clarification.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany and authorized to work as a contractor\u002Femployee as applicable; Fluent in English and German (C1+), able to write clear rationales in both; Strong analytical skills with the ability to compare nuanced outputs and justify rankings; High attention to detail and comfort following strict annotation guidelines compliance; Ability to perform repetitive evaluation and validation tasks while maintaining quality targets; Reliable internet access and ability to work remotely full-time.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience with AI data labeling, RLHF, prompt evaluation, or large language model evaluation; Familiarity with LLM behavior, common failure modes, and quality rubrics; Experience in QA evaluation, auditing, or training data quality programs; Self-driven, organized, and comfortable working independently with remote teams; Background in linguistics, translation, content moderation, or technical writing is a plus.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour, full-time remote. Final rate within range depends on skills assessment, evaluation performance, and role alignment.",{"h2":43,"desc":44},"How to Apply","Apply to Rexzone with a brief summary of your English\u002FGerman proficiency, Germany location, and any experience in evaluation, QA, data labeling, or LLM-related work. Selected candidates will complete a short qualification assessment focused on ranking, reasoning, and annotation quality.","AI Data Operations"]