We’re an AI-native venture studio building and launching web-based AI products. We need a Manual QA Tester who can test the complete product—not only the interface, but also the APIs, backend behavior, data flow, and AI-generated results.
This is a QA-first role. Full-stack engineering experience is strongly preferred because we want someone who can inspect code, diagnose likely root causes, and potentially fix smaller issues.
What you’ll do:
- Manually test AI-powered web applications across desktop and mobile browsers.
- Test complete user journeys, including registration, authentication, forms, reports, payments, permissions, and account states.
- Test frontend behavior, responsive layouts, accessibility, loading states, error messages, and overall usability.
- Test backend APIs, request and response data, authentication, validation, timeouts, rate limits, and failure handling.
- Validate data across the browser, API, and database when access is provided.
- Red-team AI and LLM functionality for hallucinations, fabricated facts, unsupported claims, prompt injection, inconsistent results, bias, and unsafe output.
- Verify that AI-generated claims, scores, sources, and recommendations are supported by evidence.
- Test low-data and ambiguous cases where the AI should admit uncertainty instead of inventing an answer.
- Create and maintain test cases, regression checklists, and launch-readiness reports.
- Report bugs with severity, reproducible steps, expected versus actual behavior, screenshots or recordings, logs, and relevant network requests.
- Retest fixes and provide a clear launch recommendation: ready, conditionally ready, or not ready.
Required qualifications
- At least 3 years of hands-on manual QA or software testing experience.
- Experience testing modern web applications from end to end.
Strong understanding of functional, exploratory, regression, integration, usability, and negative testing.
- Ability to test REST APIs using Postman, Insomnia, curl, or similar tools.
- Confidence using browser developer tools, including the console, network panel, storage, cookies, and request inspection.
- Ability to write concise bug reports that developers can reproduce without additional explanation.
- Experience testing responsive behavior across multiple browsers and screen sizes.
- Strong written English and reliable communication.
- Comfortable working independently in fast-moving, pre-launch environments.
Strongly preferred
- Direct experience testing AI, generative AI, agents, RAG systems, or LLM-powered products.
- Understanding of hallucinations, non-deterministic outputs, context limits, retrieval failures, prompt injection, and model fallbacks.
- Full-stack engineering experience with frontend and backend systems.
- Ability to read or troubleshoot HTML, CSS, JavaScript, TypeScript, React, Node.js, SQL, and serverless applications.
- Experience tracing a frontend defect through an API request to the likely backend cause.
- Ability to implement small, well-scoped fixes and submit them through Git and pull requests.
- Experience with Cloudflare, serverless functions, databases, authentication, payments, or third-party AI APIs.
- Familiarity with Playwright, Cypress, Selenium, or another automation framework.
- Accessibility testing knowledge, including WCAG.
What success looks like
- You will help us catch failures that ordinary UI testing misses: plausible-but-wrong AI answers, unsupported claims, silent backend failures, inconsistent scoring, entitlement problems, exposed credentials, broken mobile flows, and errors that appear to users as successful results.
- We care more about judgment and prioritization than the total number of bugs reported.
How to apply
- Start your proposal with the word LAUNCH to show you've read through the requirements. Submit a response to these questions, we are limited on time and need to do a deep pre-screen, followed by a short introductory video call to the candidate we select.
1. What have you personally delivered? Provide two relevant AI or LLM products you personally tested.
For each, include:
- Product or redacted project description
- Models or AI providers involved
- Frontend, backend, API, database, and mobile responsibilities
- What you personally discovered or improved
- A URL, test plan, bug report, Loom, PR, or redacted work sample (Do not send only a company name or list of responsibilities)
2. Walk us through your end-to-end process from receiving an unfamiliar pre-launch product to issuing a final go/no-go recommendation. Include a sample of the deliverables we would receive.
3. What AI failure have you personally found that normal QA would miss and how did you fix it?
4. How would you test an AI feature that produces different answers for the same input? Explain your exact methodology for determining whether it passes or fails.
5. Describe a defect you traced from the frontend through an API or backend service.
6. How would you test our complete launch flow?
7. Which tools and technologies can you test or troubleshoot?
8. . Can you make small code fixes? If yes, share your GitHub, portfolio, or a relevant example. **We will not select a candidate who can't show previous work, projects, client recommendations or references**
9. What are your hourly rate, timezone, weekly availability, and earliest start date?
If NDA prevents naming a client, anonymize the client but explain your personal contribution precisely.
**Generic proposals that do not answer these questions will not be considered.**
Required:
1. Past work/projects, Linkedin Profile, any customer references/recommendations.
2. NDA will be required before accessing confidential products.
We have multiple products in development and many more ideas in the pipeline. This will begin as a paid trial so we can evaluate your work and how well we collaborate. If successful, we plan to hire you on an ongoing, project-by-project basis.