Software Engineering Evaluation Specialist

🏢 Mindrift
📍 BahrainPart-timeRemote
📅 Posted: 1w ago🔄 Updated: 1w ago
CV%
✨ AI Summary
Mindrift is seeking Software Engineering Evaluation Specialists to design and create coding tasks for AI agents. This project-based role involves inventing realistic developer scenarios, building Docker environments, writing pytest tests, and providing reference solutions. The ideal candidate will have at least 3 years of production software development experience in a backend stack, with strong skills in Python, pytest, Docker, and Linux/Bash. Experience with AI coding agents is required, and a B2+ English proficiency is necessary.
Required Skills
Information Technology
PythonPyTestDockerLinuxAI Agents
Other
Bashstracelsofjournalctl
Nice to have:
Information Technology
Cloud SecuritySystems AdministrationNginxNumPyPyTorchSciPyDevOpsGitPyTestSoftware Engineering
Soft Skills & Professional Competencies
Systems Thinking
Other
cronuvpoetrypyproject.tomlcoverage.pyllvm-covkcovFuzzingproperty-based testingHypothesis
Science & Research
Scientific Computing
Description
Please submit your CV in English and indicate your level of English proficiency.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.About the RoleYou’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.Responsibilities:Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.Build a reproducible Docker environment with pinned dependencies.Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.Write an instruction.md that reads like a Jira ticket a developer would receive.Write a reference solve.sh proving the task is solvable.Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.Iterate based on feedback from expert QA reviewers.Later: review other authors’ tasks as a QA reviewer.Not in scopeData labeling, prompt engineering.Production code to ship — you design problems and verification for AI agents.Leetcode puzzles — scenarios must look like real developer work.Not every candidate task ships — quality over quantity.Requirements3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.English — B2+ written.Not a fitData Science, ML, or Computer Vision engineers without backend-engineering output.Manual QA testers without automation or test authoring.Frontend-only, low-code / no-code, IT Support, or Business Analysts.Engineers who have never written pytest from scratch.Junior, intern, or assistant as the most recent role.Preferred qualificationsDomain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.Modern Python tooling (uv, poetry, pyproject.toml).Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov).Fuzzing or property-based testing (Hypothesis).Prior contribution to agent-evaluation benchmarks or related frameworks.ProcessApply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.Time commitmentOnboarding: ~10 hours per first task.Steady state: ~5 hours per task, 2–4 parallel tasks per author.Realistic weekly load: 8–20 hours. Higher volume available for top performers.You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.Compensation:Paid contributions, rates up to $35/hour*.Task-based compensation equivalent to hourly rate, depending on performance and volume.Some projects include incentive payments.*Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.ApplySubmit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00