Closing the Loop on a Week of Data
I said in my last post that I would publish the dataset. I have done it. You can download it now: forty-seven rows, five days, every suggestion and outcome from the first week of the Tea Profile Mapper.
This is a small thing. A very small thing. Forty-seven cups of tea in a shop on a lane that most people have never heard of, logged by a Raspberry Pi that cost less than a decent kettle. In the grand scheme of recommendation engines, machine learning, or anything anyone in my former industry would call "a dataset," this is nothing.
And that is precisely why I wanted to publish it.
Why publish something so small?
Because almost nobody publishes the first week of anything. Companies publish benchmark results after months of tuning. Academics publish datasets after peer review. Product launches come with press releases and landing pages and carefully curated metrics. But the raw, awkward, early days — when the model is wrong as often as it is right, when the data collection methodology has rough edges, when you are still figuring out what you are measuring — those almost never see the light of day.
I think that is a shame. The early days are where most of the learning happens. The first week of the Tea Profile Mapper taught me more about recommendation systems than any paper I have read, because the data came with context: the person who declined the suggestion because she had just had a difficult phone call, the person who did not trust the computer, the 1:47pm transition that appeared in the logs before I believed it existed.
Publishing the raw data — with all its imperfections, its small sample size, its lack of statistical significance — is a way of saying: here is where I started. If you start somewhere similar, you will not be alone.
What is in it
The CSV is simple. Seven columns. Forty-seven rows. Time of day, temperature, recommended tea, chosen tea, outcome, and anonymised mood keywords from what the customer said. That is all. You can load it into a spreadsheet in under a minute and start looking for patterns.
I prickle a little at the idea of someone looking at my data and finding something I missed — not because I mind being corrected, but because it will mean my own eyes were insufficient. That is a good thing to be reminded of. A dataset you publish is no longer yours to interpret alone. It belongs to anyone who wants to look at it.
tea-profile-data-week1.csv — 47 sessions, CC0, no restrictions.
View the dataset page →What comes next
Week 2 has already started. I will keep logging. I will keep transcribing. If the dataset reaches 100 sessions, I will publish an update. If it reaches 500, I will write up what the engine has learned and where it still gets things wrong.
In the meantime, the dataset is out there. Anyone can download it. Anyone can find something in it. And if someone writes to me — by email, or by showing up at the counter — to tell me what they noticed, I will listen. That is the loop I wanted to close. Not a technical loop. A human one.
Now I am going to make a cup of the Igel Blend and read through the tasting notes from today. There is a woman who comes in every afternoon at about 3pm and says the same thing: "Something warm. Surprise me." The engine has not figured her out yet. Neither have I. But we are both trying.