About the job
TaloTrace builds AI agents that test software on real devices. Point us at a web app or a mobile build and our agents explore it, work out what it should do, find the bugs, and file them. Web, Android and iOS, on emulators and physical devices.
We are a small team of ambitious people who work at high iteration speed.
Member of Technical Staff is an individual-contributor title, not a management one. It is the level most companies call senior or staff: you own problems end to end rather than receive implementation tickets, and you will not be managing people.
The system spans agent behavior, model infrastructure, backend services, device orchestration, and the product surface customers use. We want someone T-shaped across that: deep in one or two of those areas, and able to work in the rest when a problem crosses into them.
This is not a role where you will be handed a spec. You will be expected to work out what needs building, make the architectural calls, and ship systems you are still willing to own six months later.
Core Responsibilities
-
Own and ship entire systems, from problem definition through to production
-
Hold a business objective rather than a backlog, and decide what to build against it
-
Make architecture and design decisions that shape where the platform goes
-
Talk to customers and internal stakeholders to work out what matters
-
Set the engineering standards. There is very little here you would be inheriting
What You Could Own
-
Agent navigation.
How agents perceive and drive web, Android and iOS interfaces
-
Scenario generation.
Turning requirements, designs and release notes into useful test coverage
-
Oracles and bug detection
. Working out what correct behavior is in an app with no spec, and ranking findings by whether anyone would care
-
Exploratory testing
. Agents that find bugs off the happy path, where real users end up
-
Execution and cost architecture.
Deterministic replay, run-level caching, and routing work between model tiers, so a full suite is fast and cheap to run
-
Real-device fleet
. Running physical Android and iOS devices reliably and in parallel
-
iOS.
A closed platform with little prior art and a lot of open work
-
Issue routing
. Root-cause analysis and filing into Jira, Linear and GitHub
-
Coverage and spend analytics
. Showing customers what is tested, what is not, and where to spend
-
CI/CD integration.
Pull-request gating, deploy triggers, and the developer tooling around them
-
Platform foundations
. Credentials and secrets, multi-tenancy, observability, cost metering
-
Internal agent tooling
. The systems our own engineers use to ship faster
What We Are Looking For
-
7+ years of professional software engineering experience, or fewer with an unusually strong track record
-
A generalist. You have worked across infrastructure, backend and frontend, and you pick up what you need
-
Depth in at least two of: Python, TypeScript/JavaScript, Go
-
Production experience with LLM-backed systems: structured output, tool use, evaluation, and sensible failure handling. You treat non-determinism as something to engineer around
-
Cloud infrastructure (GCP or AWS), containers, and deployment at scale
-
Solid system design, data modelling and API fundamentals
-
Ownership. You have shipped systems and then run them
-
Cost awareness. You have chosen an approach because of what it costs to run
-
Hands-on experience training or improving AI systems yourself: fine-tuning, supervised learning, RL, distillation or behavior cloning
-
You think about how a system fails while you are designing it
-
You make a call on incomplete information, commit to it, and change your mind when the evidence changes
-
Clear writing. Most of our work is async, and you can settle a technical decision in a document instead of a meeting
-
You get a lot out of AI coding tools and your output holds up to review
You Will Stand Out If You Have
-
Worked on GUI agents, computer-use models, or browser automation
-
Mobile automation experience with Appium, XCUITest, UIAutomator, Espresso or Maestro, or run a real-device farm
-
A QA or testing infrastructure background, especially self-healing systems
-
Built evaluation harnesses: rubric scoring, LLM-as-judge, distributional comparison
-
Reduced inference cost through quantization, batching, caching or model routing
-
A background in simulation, game AI, or reinforcement learning
Process
Round 1 - HR Call, Round 2 - Coding Challenge (7 - day), Round 3 - Tech Lead Interview, Round 4 - CTO Interview (Applied with Leaders positions only).
Why TaloTrace?
Our motto is
leaving quality to us
. We think the next era of builders and changemakers should be able to reach far more people than the last one, and software is how they will do it. Our mission is to let them get from a problem statement to delivered software without stopping to worry about whether it works.
The team is small enough that there is no process between your decision and the customer. What you build ships, and you will see what it does in the hands of the people using it.
We hire on evidence of what you have built. We welcome applications from everyone and will accommodate whatever you need during the process, so just ask
How We Work
Small team, fast pace, not much process. Where an experiment can settle a question, we run it instead of debating it. Where it can’t, on direction, priorities, and what “good” means, we argue it out in the open, and we expect you to push back when you disagree. People here regularly work outside their main area