24-MAG is a next-gen management consulting partner built to help B2B organizations move faster and operate smarter.
QA Test Engineer
Location
New York
Posted
6 days ago
Salary
$55 - $85 / hour
Seniority
Junior
Job Description
QA Test Engineer
24-MAG
• Create comprehensive test cases confirming that benchmark tasks function as intended • Design positive, negative, boundary, and edge-case tests • Validate task requirements, expected outputs, reference solutions, and grading logic • Identify scenarios that may produce incorrect or misleading evaluation results • Ensure tests measure the intended technical capability accurately • Review complex multi-step tasks and reference solutions before finalisation • Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteria • Run tasks independently to confirm reproducibility and expected behaviour • Assess whether grading standards are clear, fair, and technically defensible • Provide actionable feedback to task authors and researchers • Investigate failures across Python scripts, test harnesses, repositories, and task environments • Diagnose unexpected behaviour within unfamiliar codebases • Reproduce reported issues and isolate their underlying causes • Correct or document environment, dependency, logic, and validation problems • Use Git-based workflows to support structured review and collaboration • Develop practical checklists and repeatable review procedures for benchmark quality • Improve consistency across task validation, testing, and approval workflows • Document findings clearly so authors can resolve issues efficiently • Track recurring defects and recommend preventive quality measures • Collaborate closely with researchers, task authors, and other technical reviewers • Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses • Identify cases where models can receive credit without completing the intended reasoning or technical work • Test whether benchmark tasks remain robust across alternative approaches • Strengthen evaluation criteria to maintain reliable and meaningful benchmark scores • Distinguish valid solution diversity from unintended task exploitation
Job Requirements
- At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical role
- Demonstrated experience designing test cases and quality-review processes
- Strong end-to-end debugging skills across complex technical systems
- Working proficiency in Python and Git
- Comfort navigating unfamiliar codebases, repositories, and execution environments
- Exceptional attention to detail and strong written documentation habits
- Ability to identify ambiguity, edge cases, hidden assumptions, and quality gaps
- Capacity to work independently through open-ended technical problems
- Reliable availability for approximately 35 hours per week
Related Guides
Related Categories
Related Job Pages
More SDET Jobs
• Lead the design, development, and maintenance of automated testing frameworks and quality assurance strategies across web, mobile (Android/iOS), and API-based applications. • Ensure the quality, reliability, and scalability of products by driving automation initiatives and improving test coverage. • Design and implement automated test solutions and define corrective and preventive actions based on customer requirements and business needs. • Build and enhance automation frameworks and integrate automated testing into CI/CD pipelines. • Own the end-to-end testing lifecycle and take responsibility for the quality of products delivered to customers. • Validate complex business logic, calculations, and mathematical operations to ensure data accuracy and functional correctness. • Perform data validation and reconciliation testing using SQL queries, database analysis, Excel, VLOOKUPs, and other analytical tools. • Test and validate RESTful APIs, ensuring data integrity and service reliability.
Senior QA Automation Engineer
Buildout, Inc. Buildout, Inc. offers an end-to-end solution for marketing commercial real estate listings and empowers brokerages nationally to showcase their brand and gr
• Own the E2E test suite. Design, build, and maintain end-to-end automated tests in Playwright (TypeScript) covering our critical user journeys across products. • Backfill and prioritize coverage. Audit what we have, identify the highest-risk gaps, and systematically close them. • Kill flakiness. Own suite reliability end to end: root-cause flaky tests, establish quarantine and retry policies, manage test data and environment stability, and keep signal-to-noise high enough that a red build means something. • Shape CI integration. Partner with our DevOps engineer on where tests run in the pipeline, execution frequency, parallelization, runtime budgets, and merge/deploy gating. • Test at the right layer. Know when a check belongs at the API or integration layer instead of the browser, and structure the suite so it stays fast as it grows. • Enable the team. Coach developers on writing good Playwright tests, review test code, and establish patterns, utilities, and documentation so test authorship scales beyond you. • Report on quality. Co-own quality metrics including test coverage of critical paths, suite reliability, and defect escape rate, and communicate trends to engineering leadership. • Use modern tooling. Leverage AI-assisted development (we use Claude Code across engineering) and Playwright's evolving tooling to accelerate test authoring and maintenance.
Python Software Engineer
24-MAG24-MAG is a next-gen management consulting partner built to help B2B organizations move faster and operate smarter.
Role Description We are sharing a specialised full-time consulting opportunity for experienced software engineers with strong Python development, debugging, version-control, technical documentation, and AI-assisted coding experience. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will design, implement, and review realistic multi-step software engineering tasks that test the capabilities of AI coding agents across Python development, environment setup, tooling, debugging, and technical problem-solving. Key Responsibilities - Software Engineering Task Design - Create realistic, multi-step software engineering challenges based on practical development workflows. - Design technically demanding problems that require implementation, debugging, environment configuration, and analytical reasoning. - Define clear requirements, constraints, expected outputs, and acceptance criteria. - Ensure tasks assess genuine software engineering capability rather than superficial code generation. - Reference Solution Development - Build complete and verifiable reference solutions in Python. - Create the supporting setup, dependencies, tests, and validation checks required for each task. - Write clean, readable, and maintainable code. - Confirm that solutions run reliably within the intended technical environment. - Document implementation decisions and expected behaviour clearly. - AI Coding Agent Evaluation - Use AI coding assistants and agent-based development tools within practical engineering workflows. - Evaluate how frontier models approach complex coding and debugging tasks. - Identify implementation errors, unsupported assumptions, inefficient approaches, and incomplete solutions. - Analyse where AI agents succeed, struggle, or exploit unintended shortcuts. - Document failure patterns and provide evidence supporting evaluation conclusions. - Peer Review & Task Refinement - Review tasks and reference solutions created by other software engineering specialists. - Assess clarity, correctness, difficulty, reproducibility, and technical fairness. - Identify ambiguous instructions, hidden assumptions, grading gaps, and environment issues. - Provide actionable feedback that improves task quality and benchmark reliability. - Collaborate closely with researchers and fellow task authors. Qualifications - At least 1 year of experience in software engineering, research engineering, or a related coding-intensive role. - Strong hands-on Python scripting, implementation, and debugging skills. - Experience developing clean, readable, and maintainable software. - Everyday fluency with Git, IDEs, repositories, and standard software development workflows. - Comfort configuring environments, dependencies, tooling, and validation processes. - Strong technical writing and documentation skills. - Ability to work independently through ambiguous and open-ended engineering problems. - Reliable availability for approximately 35 hours per week. Requirements - An MSc or PhD in computer science, software engineering, another STEM discipline, or a related technical field is highly relevant. - Equivalent practical experience in a research-intensive or engineering-intensive role may also be considered. - Academic or professional work involving significant coding, data analysis, or technical experimentation may strengthen an application. - Open-source contributions, technical projects, publications, or substantial software development work may also be valuable. Nice to Have - Experience using AI coding assistants, prompt engineering methods, or agent-based workflows. - Previous work in AI training, model evaluation, or benchmark development. - Background authoring technical tasks, reference solutions, or grading criteria. - Familiarity with automated testing, CI/CD workflows, containers, or reproducible environments. - Experience reviewing code or technical assignments created by other engineers. - Knowledge of agentic AI systems and multi-step coding evaluations. - Strong ability to identify edge cases, unintended shortcuts, and subtle implementation issues. - Experience collaborating with AI research or evaluation teams. Why This Opportunity - Apply practical software engineering expertise to frontier AI evaluation. - Design realistic coding tasks grounded in professional development workflows. - Help researchers understand where advanced AI coding agents succeed and fail. - Work across Python implementation, debugging, environment setup, and benchmark development. - Collaborate closely with AI researchers and experienced software engineers. - Participate in a structured full-time remote role with competitive hourly compensation. Contract Details - Full-time W-2 contingent employment opportunity. - Fully remote within the United States. - Expected commitment of approximately 35 hours per week. - Competitive rates between $55–$85 per hour depending on expertise and project scope. - Individual tasks may require one to two days of focused engineering work. - Work may include task design, Python development, reference-solution creation, AI agent evaluation, peer review, and technical documentation. - Engagement scope and duration may evolve according to project requirements and performance.
Robotics Software Engineer, C++ and Python
Simbe RoboticsFounded in 2014, Simbe builds automation solutions for retailers. The company's first product, Tally, is the world's first entirely autonomous shelf-auditing an
Robotics Software Engineer (C++ & Python) San Francisco Bay Area Product Engineering – Robotics Engineering / Full-time / Hybrid Simbe Robotics is a leading retail robotics company providing in-store intelligence solutions that help retailers optimize operations, improve shelf execution, and deliver valuable data insights. Our autonomous robots and multi-modal data collection systems are transforming how retailers manage inventory and make data-driven decisions. Simbe is looking for a strong Python & C++ engineer. In this role, you will be working with our robot software engineering team on the code that drives our Tally(TM) autonomous robots. You will work on all aspects of the Tally stack including but not limited to navigation, perception, autonomous behaviors, hardware drivers, cloud integration, and infrastructure management. Your primary objective will be to build, maintain, and evolve the Tally software stack to make our robots better, faster, smarter, easier, and bulletproof to failure. Responsibilities - Maintaining and extending the Tally software stack - Working on and developing new software packages to be shared across Simbe software teams - Improving Tally's autonomy, navigation, perception, and human-robot interaction (HRI) behaviors. - Assist in our ongoing Devops & CI/CD development - Evaluating third-party SW (ROS, etc.) packages for integration into our stack Qualifications - BS, MS, or PhD in Computer Science or related field highly recommended but not required - ~ 2 years of industry experience - Extremely adept in both C++ and Python programming - Proficient in shell scripting, preferably with Bash and Python - Good understanding of the Robot Operating System (ROS) and core concepts such as nodes, messages, topics, services, parameters, build system, etc. Understanding of both ROS1 and ROS2 is strongly preferred - Experience writing ROS nodes - Well-versed in source control systems, particularly Git - Experience working with Ubuntu or other Debian-based Linux distributions - Familiarity with modern software development methodologies (e.g. continuous integration/deployment, scrum, automated regression testing) - Experience in packaging and deploying software in production environments Recommended Qualifications - Experience with databases, especially redis - Familiarity with Docker containers - Experience with Nvidia Jetson platform - Experience with cloud computing platforms (GCP, AWS, Azure, etc) - Experience managing large numbers of connected IoT devices (e.g. robots, wearables, phones, smart home) $100,000 - $160,000 a year The base salary offered is based on market location, and may vary further depending on individualized factors for job candidates, such as job-related knowledge, skills, experience, and other objective business considerations. Subject to those same considerations, the total compensation package for this position may also include other elements, including equity compensation, in addition to a full range of medical, financial, and/or other benefits. Simbe’s approach emphasizes total rewards - base pay, equity, incentives, and benefits - rather than viewing compensation as cash alone. We believe the full package, including ownership through equity and well-being support, is what drives engagement, retention, and alignment with our mission. Simbe Values: R. E. T. A. I. L. Result Driven - We are customer-centric and results-driven. We strive to create immense value for our team, partners, customers, and investors. Empathetic - We are sensitive and mindful. We support each other in challenging times, both professionally and personally. Transparent - We highly value open communication internally, and with our partners and customers. We are receptive to feedback. Agile - We are agile and always eager to learn. We quickly adapt to changes and customer needs Innovative - We are bold and innovative, with an intense focus on product design and user experience. Leaders - We strive for excellence. We are accountable, the best at what we do, and leaders in our field. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


