[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-llm-evaluator-jobs-berlin-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. The role is remote, but you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will evaluate, rank, and QA model-generated outputs, write rationales, perform validation checks, and follow annotation guidelines compliance to improve training data quality and model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We provide onboarding and rubrics; strong analytical skills, careful reasoning, and attention to detail are essential.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required for bilingual evaluation and writing tasks.","What languages are required?",{"A":20,"Q":21},"Tasks can span general knowledge, writing quality, instruction following, reasoning, factuality, and content safety labeling within large language model evaluation workflows.","What domains are covered?","llm-evaluator-jobs-berlin-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support AI\u002FLLM workflows through RLHF, large language model evaluation, and training data quality improvements by assessing, ranking, and validating model outputs with clear rationales and annotation guidelines compliance.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate and improve large language model outputs across real-world tasks. You will apply RLHF-style preference ranking, prompt evaluation, and QA evaluation to drive training data quality and model performance improvement. You will follow annotation guidelines compliance requirements, document decisions with strong reasoning, and help maintain consistent, safe, and reliable training datasets.",{"h2":31,"desc":32},"Key Responsibilities","Evaluate and rank model-generated outputs in English and German using RLHF-style preference signals. Perform LLM evaluation and prompt evaluation against defined rubrics, including reasoning quality, completeness, and factuality checks. Write concise rationales that explain ranking decisions and support large language model evaluation. Execute QA evaluation and validation audits to ensure training data quality and annotation guidelines compliance. Identify edge cases, ambiguity, and policy risks; perform content safety labeling when required. Apply data labeling standards consistently and escalate unclear cases with proposed guideline clarifications. Track errors and patterns, propose improvements, and contribute to model performance improvement through feedback loops. Maintain high throughput while preserving attention to detail and consistent judgment across tasks.",{"h2":34,"desc":35},"Basic Qualifications","Based in Germany and authorized to work remotely from Germany. Fluent in both English and German (reading, writing, and comprehension). Strong analytical skills with the ability to compare alternatives and justify decisions using clear reasoning. High attention to detail and ability to follow annotation guidelines compliance standards consistently. Comfortable working with web-based labeling tools and structured evaluation rubrics.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in AI data labeling, LLM evaluation, RLHF, prompt evaluation, QA evaluation, or content moderation workflows. Familiarity with common LLM failure modes (hallucinations, instruction-following issues, safety\u002Fpolicy violations). Self-driven and reliable in a fully remote environment, with strong time management and ownership. Experience writing clear rationales and applying consistent evaluation criteria under ambiguity.",{"h2":40,"desc":41},"Compensation and Work Setup","This is a full-time, remote role for candidates based in Germany. Compensation is $35–$40 USD per hour, depending on assessment performance and task alignment. You will receive onboarding materials, evaluation rubrics, and ongoing calibration to support consistent training data quality and model performance improvement.",{"h2":43,"desc":44},"How to Apply","Apply to Rexzone with your resume\u002FCV and a brief note confirming you are based in Germany and fluent in English and German. If selected, you will complete a short qualification assessment focused on large language model evaluation, ranking, and annotation guidelines compliance.","AI Data Operations"]