[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-fullstack-repo-code-annotator-remote":3},{"Slug":4,"Header":5,"Ques":39,"job_category":72},"fullstack-repo-code-annotator-remote",{"title":6,"desc":7,"content":8},"Repo-based Code Annotator","Remote opportunity for experienced engineers to build reproducible Docker-based test environments, strengthen unit test coverage, validate SWE-Bench\u002FTerminal-Bench workflows, and write clear, standardized task documentation. Compensation: USD $80–$120 per day, based on skills and experience.",[9,12,15,18,21,24,27,30,33,36],{"h2":10,"desc":11},"About the Role","As a Repo-based Code Annotator, you will design and maintain reproducible, standardized test environments to replicate known issues or produce expected outputs according to defined procedures. You will review and improve unit test coverage, validate the completeness and rationality of test sets, and ensure that workflows for tasks related to SWE-Bench and Terminal-Bench are precisely aligned. Documentation quality and reproducibility are central to the role.",{"h2":13,"desc":14},"Key Responsibilities","- Build deterministic Docker images and environments to reproduce issues and generate expected outputs.\n- Review, extend, and refactor unit tests to evaluate correctness, stability, and coverage of target code.\n- Validate test set completeness and rationality; align task workflows with SWE-Bench and Terminal-Bench requirements.\n- Write clear, standardized documentation (e.g., task.yaml and README) to ensure consistency and reproducibility.\n- Develop Python-based task harnesses, automation scripts, and utilities supporting testing workflows.\n- Maintain clean Git\u002FGitHub practices, producing high-quality, reproducible pull requests.",{"h2":16,"desc":17},"Required Skills","- Strong proficiency with Linux command line and Shell scripting; comfort with tools such as grep, sed, awk, curl, and jq.\n- Expert-level Python for building task harnesses, writing unit tests, and automating workflows.\n- Solid Docker experience, including authoring Dockerfiles and building reproducible environments.\n- Familiarity with pytest (or similar), with techniques for mocking data and controlling randomness.\n- Competence with Git\u002FGitHub workflows for collaborative, reproducible development.",{"h2":19,"desc":20},"Professional Background","- Degree or equivalent experience in Computer Science, Software Engineering, Artificial Intelligence, or related fields.\n- Relevant experience in Software Development, Test Engineering, DevOps, or Data Engineering.\n- Preference for contributors to open-source projects, especially in automated testing, CI\u002FCD, and containerization.",{"h2":22,"desc":23},"Bonus Points","- Proficiency in Go or Rust for performance-critical tooling.\n- Familiarity with Docker Compose or Podman and other sandbox technologies.\n- Ability to design datasets\u002Ftasks that mitigate task cheating.\n- Understanding of scientific benchmark design principles: fairness, repeatability, scalability.\n- Experience with automated testing systems or CI\u002FCD, and a cross-disciplinary perspective.",{"h2":25,"desc":26},"Compensation","USD $80–$120 per day, dependent on actual skills and experience. The rate reflects the role's emphasis on reproducibility, robust testing, precise workflow alignment, and comprehensive documentation.",{"h2":28,"desc":29},"Work Setup & Collaboration","- Fully remote, suitable for distributed teams and asynchronous collaboration.\n- Collaboration through GitHub issues, pull requests, and code reviews.\n- Documentation-first approach with standardized procedures and reproducible outcomes.\n- Emphasis on deterministic builds, clear test artifacts, and traceable changes.",{"h2":31,"desc":32},"Tools & Technologies","- Linux, Shell scripting (grep, sed, awk, curl, jq).\n- Python, pytest, and testing utilities for mocking and randomness control.\n- Docker\u002FDockerfiles; familiarity with Docker Compose\u002FPodman is a plus.\n- Git\u002FGitHub, CI\u002FCD systems.\n- SWE-Bench and Terminal-Bench task workflows and validation.",{"h2":34,"desc":35},"Success Indicators","- Deterministic, reproducible builds and environments.\n- Measurable improvements in unit test coverage and stability.\n- Clear, actionable task.yaml\u002FREADME documentation that enables consistent execution.\n- Validated test sets and workflows aligned with SWE-Bench\u002FTerminal-Bench guidelines.",{"h2":37,"desc":38},"Location & Schedule","- Remote role with flexibility across time zones.\n- Output-focused collaboration; occasional overlap for reviews or syncs may be requested.",{"title":40,"content":41},"Frequently Asked Questions",[42,45,48,51,54,57,60,63,66,69],{"Q":43,"A":44},"What does a Repo-based Code Annotator do day-to-day?","You will build reproducible Docker environments to replicate issues or expected outputs, create and refine unit tests, validate SWE-Bench\u002FTerminal-Bench task workflows, write task.yaml\u002FREADME documentation, and implement Python-based harnesses and automation with reproducible Git\u002FGitHub practices.",{"Q":46,"A":47},"How strong should my Python and Docker skills be?","You should be comfortable authoring Dockerfiles, building deterministic images, and writing Python test harnesses and automation. Familiarity with pytest, mocking, and controlling randomness is expected.",{"Q":49,"A":50},"Is this position fully remote?","Yes. The role is fully remote and suited to asynchronous collaboration across different time zones.",{"Q":52,"A":53},"What is the compensation range?","USD $80–$120 per day, based on demonstrated skills and relevant experience.",{"Q":55,"A":56},"Which tools are used most frequently?","Linux CLI tools (grep, sed, awk, curl, jq), Python, pytest, Docker, Git\u002FGitHub, and CI\u002FCD systems. Familiarity with SWE-Bench and Terminal-Bench workflows is beneficial.",{"Q":58,"A":59},"Are open-source contributions required?","They are not required but are preferred—especially contributions in automated testing, CI\u002FCD, or containerization—as they demonstrate strong reproducibility and code quality practices.",{"Q":61,"A":62},"How is success measured in this role?","Success includes deterministic builds, improved test coverage and stability, validated and rational test sets, benchmark-aligned workflows, and high-quality documentation that enables reproducible execution.",{"Q":64,"A":65},"Are Go or Rust needed for this role?","They are not required, but proficiency in Go or Rust is a plus for performance-focused utilities.",{"Q":67,"A":68},"Will I design tasks or datasets?","Yes, you may. Designs should consider preventing task cheating and follow benchmark principles such as fairness, repeatability, and scalability.",{"Q":70,"A":71},"What Git\u002FGitHub workflow is expected?","Use standard branching, clear commits, thorough tests, and documentation. Pull requests should be reproducible and easy to review.","Code Annotation"]