Data Engineer — Visibility Dataset
Build the longitudinal dataset of what AI recommends — millions of immutable observations across ten assistants, six languages and every tracked market.
About the role
Growli's moat is not a feature — it is a dataset nobody else is collecting: how every major AI assistant answers real buying questions, week after week, across languages and cities. Today that dataset lives in a pipeline built for speed of iteration; your job is to give it the foundations of the category-defining asset it needs to become. You will own storage, modeling, quality and access: append-only observation stores, aggregation layers that serve dashboards in milliseconds, quality monitors that catch drift before customers do, and the internal analytics that tell us which markets we should measure next. Every future product — benchmarks, industry indexes, the national business repertoire — gets built on what you lay down here.
What you will do
- Own the observation store: immutable prompt-run data, content-addressed answers, and the schema evolution strategy around them.
- Build aggregation layers so per-assistant, per-city and per-competitor views stay fast as data grows a hundredfold.
- Design data-quality monitoring: coverage, freshness, detection drift and cost per observation — alarmed, not eyeballed.
- Make the dataset queryable for product experiments without risking the production hot path.
- Model cross-business benchmarks that respect tenant isolation and our anonymity floors.
- Own the migration path as we outgrow the current storage engine — measured, reversible, boring.
- Partner with the AI engineer on labeled evaluation sets and with product on new evidence types.
What success looks like
- Dashboard queries stay under 100ms while observations grow by an order of magnitude.
- Data-quality incidents are caught by monitors, not by customers, with a runbook for each class.
- An industry benchmark built purely from the dataset ships as a customer-facing feature.
- Storage cost per observation goes down while durability guarantees go up.
What we are looking for
- 4+ years building data systems in production — pipelines, warehouses or high-volume OLTP.
- Strong SQL and Python; you model data for the queries it must answer, not for the diagram.
- You have lived through a storage migration and can tell us what you would never do again.
- You treat data quality as a product with SLOs, not a cleanup chore.
- Pragmatic about infrastructure: the boring proven tool beats the exciting new one.
- Bonus: experience with search, analytics products, or multi-tenant SaaS datasets.
How to apply
Email us your CV and a short note on why this role. No cover letter template — tell us what you would own in the first three months.
Apply by emailApplications are read by the founder. We reply either way.
Posted September 1, 2026
