[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-bilingual-ai-evaluator-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a full-time remote role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as evaluating outputs, ranking responses for RLHF, writing reasoning-based rationales, completing QA evaluation, and validating training data quality against annotation guidelines.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. If you can follow rubrics, apply annotation guidelines compliance, and provide consistent evaluations with clear reasoning, you can succeed. Prior data labeling or evaluation experience is a plus.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both German and English is required, as you will evaluate and compare content in both languages.","What languages are required?",{"A":20,"Q":21},"You may evaluate general knowledge, customer support-style prompts, writing quality, reasoning, safety and policy adherence, and other everyday use cases relevant to prompt evaluation, content safety labeling, and model performance improvement.","What domains are covered?","bilingual-ai-evaluator-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based AI Generalist Trainers to support large language model evaluation through RLHF, prompt evaluation, and QA evaluation—improving training data quality and driving model performance improvement via consistent ranking, validation, and annotation guidelines compliance.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate model-generated outputs across common user scenarios and help improve AI\u002FLLM workflows. Your work will focus on RLHF-style preference ranking, large language model evaluation, and training data quality initiatives. You will apply annotation guidelines compliance, write clear rationales, and conduct QA evaluation to support model performance improvement. This is a full-time remote role requiring bilingual fluency in German and English.",{"h2":31,"desc":32},"Responsibilities","Evaluate and compare model-generated responses in German and English using defined rubrics; Rank outputs for RLHF datasets and document reasoning\u002Frationales for preference decisions; Perform prompt evaluation and QA evaluation to identify inconsistencies, factual errors, or policy violations; Validate labeled data for training data quality, including edge-case handling and guideline adherence; Apply annotation guidelines compliance and escalate ambiguous cases with well-structured notes; Conduct content safety labeling and ensure safe-completion standards are met; Review peer work, provide calibrated feedback, and support ongoing quality improvement; Track recurring error patterns and propose rubric or guideline clarifications for model performance improvement.",{"h2":34,"desc":35},"Basic Qualifications","Based in Germany and authorized to work from Germany; Fluent in German and English (written and reading comprehension required for detailed evaluations); Strong analytical skills with the ability to weigh tradeoffs and justify rankings with clear reasoning; High attention to detail and consistency when following rubrics and annotation guidelines compliance; Comfortable working with ambiguous tasks and applying policy\u002Frubric judgment; Reliable internet connection and ability to meet quality and throughput targets in a remote setting.",{"h2":37,"desc":38},"Preferred Qualifications","Experience with data labeling, content safety labeling, or QA evaluation in AI\u002FML programs; Familiarity with LLM evaluation concepts (e.g., RLHF, preference ranking, prompt evaluation); Background in linguistics, translation, writing, research, or quality assurance; Self-driven, organized, and comfortable owning tasks end-to-end with minimal supervision; Experience working with annotation tools and structured feedback loops for training data quality.",{"h2":40,"desc":41},"Pay And Employment Details","Full-time, remote. Compensation is $35–$40 USD per hour depending on assessment results and role alignment. You will contribute directly to large language model evaluation, training data quality, and model performance improvement for Rexzone clients.",{"h2":43,"desc":44},"How to Apply","Apply with an up-to-date resume\u002FCV and a brief note describing your experience evaluating written content in German and English. Highlight any work involving data labeling, QA evaluation, prompt evaluation, or annotation guidelines compliance. Rexzone reviews applications on a rolling basis.","AI Data Operations"]