AI Engineer — Answer Intelligence
Own the pipeline that asks thousands of real buyer questions to ChatGPT, Claude, Gemini and seven other assistants — and turns their answers into evidence businesses act on.
About the role
Growli's entire product stands on one dataset: real assistant answers to real buying questions, collected continuously across ten AI platforms and six languages. You will own that collection and analysis pipeline end to end — probe orchestration, mention and ranking detection, source extraction, change detection — and push its accuracy from good heuristics toward measured, benchmarked truth. This is applied AI engineering with an unusual property: every improvement you ship is immediately visible to customers as better evidence. As models change monthly, the pipeline that measures them has to evolve just as fast — you will be the person who keeps us ahead of that curve, and your decisions will define how the whole category measures AI visibility.
What you will do
- Own the multi-assistant probe pipeline: scheduling, resilience, cost control and per-provider throttling across ten platforms.
- Improve mention, position and sentiment detection across six languages — measured against labeled answer sets, not vibes.
- Build the source-extraction layer that traces which pages and domains shape each assistant's answers.
- Design evaluation harnesses that catch detection regressions before customers ever see them.
- Track model releases and provider API changes, and adapt the pipeline within days, not quarters.
- Work with the product team to turn raw answer data into new measurable capabilities — geo visibility, competitor share, change alerts.
- Keep the longitudinal dataset immutable, well-modeled and cheap to query as it grows by millions of observations.
What success looks like
- Detection accuracy has a number, a benchmark suite, and a visible upward trend.
- A new assistant or model can be added to the probe fleet in under a day.
- Provider incidents degrade gracefully instead of producing silent data gaps.
- At least one new evidence type customers pay for exists because you built it.
What we are looking for
- 3+ years of production engineering with Python; you have shipped and owned data or ML pipelines.
- Hands-on experience with LLM APIs, their failure modes, and prompt-level evaluation.
- You treat measurement as an engineering problem: baselines, regressions, labeled data.
- Comfortable owning a system end to end — architecture, cost, monitoring, incident response.
- You write clearly; a design doc from you is shorter and sharper than the meeting it replaces.
- Bonus: NLP experience across multiple languages, or prior work on search or SEO tooling.
How to apply
Email us your CV and a short note on why this role. No cover letter template — tell us what you would own in the first three months.
Apply by emailApplications are read by the founder. We reply either way.
Posted September 1, 2026
