Tea Profile Dataset — Week 1
When I said in my last post that I would publish the first week of data from the Tea Profile Mapper, I meant it. It took me the rest of the day to format everything properly — not because the data is complicated, but because I wanted to check every row. Forty-seven cups of tea. Forty-seven entries in the log. Forty-seven chances to get the transcription right.
I hm-sniffed over the spreadsheet for longer than I should have. There is something about seeing data in a grid — time, temperature, outcome — that changes the way you feel about it. When it was a notebook with handwriting, it felt like notes. When it became a CSV, it felt like evidence.
6 tea types recommended · Temperatures: 88°C–97°C
What the data contains
Each row represents one customer interaction at the counter. The columns are:
- session — sequential identifier (1–47)
- date — the day
- time — when the customer walked in
- temperature_c — kettle temperature at time of brewing
- recommended_tea — what the engine suggested
- chosen_tea — what the customer actually ordered (may differ)
- outcome — accepted / modified / declined
- mood_words — anonymised keywords from what the customer said about their state
I have removed any identifying details. The mood words are selected phrases, not full sentences — enough to capture the signal, not enough to reconstruct a conversation. Temperature readings are accurate to within ±0.5°C (the probe was calibrated against a known reference at 90°C before the experiment began).
The file
That is it. Forty-seven rows. No licence restrictions — it is CC0. Use it to learn, to test, to argue with, or to prove me wrong. I would be genuinely interested to hear from anyone who finds something in it that I missed.
What I hope happens
I said in my previous post that publishing the dataset is a way of closing the loop — of saying the idea produced real interactions, and here are the numbers. That is still true. But I also hope someone else looks at this and sees something I did not. A temperature pattern that correlates to something I have not considered. A particular combination of mood words that clusters differently from how I grouped them. A mistake in my methodology that I have been too close to the data to spot.
That is the point of publishing raw data: it lets other people see through their own eyes, not through your summary. I snuffle at this CSV like I snuffled at the notebook, but I have been looking at it for a week. Fresh eyes would be welcome.
Week 2 starts tomorrow. The kettle will be on at 9am. The Raspberry Pi will be humming. And if the data from week 1 tells me anything useful, it is that the 1:47pm transition to the Igel Blend is real — or at least, it was real for five days in June. I will find out whether it holds.