[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-data-labeling-remote-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":48},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a fully remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will evaluate and rank model-generated outputs, perform QA evaluation for annotation guidelines compliance, validate bilingual content, and write reasoning-based rationales to improve training data quality and support model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. You will be trained on evaluation rubrics, RLHF-style ranking, and large language model evaluation processes.","Do I need AI experience?",{"A":17,"Q":18},"You must be fluent in German and English, with strong reading and writing skills in both languages.","What languages are required?",{"A":20,"Q":21},"You may evaluate general knowledge, customer-style questions, writing quality, reasoning tasks, and content safety labeling scenarios, depending on project needs.","What domains are covered?","ai-data-labeling-remote-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support RLHF and large language model evaluation workflows by assessing, ranking, and validating model outputs to drive training data quality and model performance improvement in a fully remote, full-time role.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42,45],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate and improve AI systems by reviewing model-generated responses, ranking alternatives, and writing clear rationales aligned with annotation guidelines. Your work supports RLHF, large language model evaluation, and training data quality initiatives that directly contribute to model performance improvement. This role requires native-level fluency in German and strong professional English, as you will work across bilingual prompts, outputs, and evaluation rubrics.",{"h2":31,"desc":32},"What You Will Do","You will conduct large language model evaluation by comparing outputs for correctness, helpfulness, safety, tone, and reasoning quality; perform RLHF-style ranking and preference judgments; execute QA evaluation to ensure annotation guidelines compliance; validate edge cases and escalate policy or ambiguity issues; and document concise, defensible rationales that can be used to improve training data quality and support model performance improvement.",{"h2":34,"desc":35},"Responsibilities","Evaluate and score model outputs against rubrics for accuracy, completeness, safety, and reasoning.\nRank multiple responses using RLHF-style preference judgments and justification.\nPerform QA checks, spot inconsistencies, and enforce annotation guidelines compliance.\nValidate bilingual (German\u002FEnglish) prompts and responses for linguistic quality and intent alignment.\nWrite clear rationales that explain evaluation decisions and support model performance improvement.\nIdentify failure patterns, propose guideline clarifications, and flag content safety risks.\nCollaborate asynchronously with operations and QA to meet throughput and quality targets.",{"h2":37,"desc":38},"Basic Qualifications","Based in Germany and authorized to work from Germany.\nFluent in German and English (reading\u002Fwriting at a high professional level).\nStrong analytical skills with the ability to evaluate nuanced reasoning and factuality.\nExcellent attention to detail and consistency in following rubrics.\nComfortable working with structured guidelines, feedback loops, and quality audits.\nReliable internet connection and ability to work independently in a remote environment.",{"h2":40,"desc":41},"Preferred Qualifications","Prior experience in data labeling, prompt evaluation, QA evaluation, or annotation.\nFamiliarity with LLM evaluation, RLHF, or model training pipelines.\nExperience applying content safety labeling or policy-based decisions.\nSelf-driven, able to manage time effectively, and proactive in clarifying ambiguous cases.\nBackground in linguistics, journalism, research, customer support, or technical writing is a plus.",{"h2":43,"desc":44},"Compensation","USD $35–$40 per hour (based on experience and assessment performance). Full-time, remote.",{"h2":46,"desc":47},"How to Apply","Apply through Rexzone with your resume\u002FCV and a brief note highlighting bilingual (German\u002FEnglish) writing experience, analytical evaluation work, and any exposure to AI\u002FLLM workflows. If selected, you will complete an evaluation aligned to large language model evaluation and annotation guidelines.","AI Data Operations"]