
Will curiosity kill the cat? From data swamps to data lakes
About
"Data is the new oil of the digital economy." That may be so, but in order to get something from this oil, it has to be extracted, refined and delivered to the recipients. Every company at some stage has collected data... and collected ... and collected ... because... you know: Big Data. But often their vision of a beautiful crystal clear data lake ends up more like a murky pond, if not a swamp. If you try to navigate the swamp or find something in the depths, you'll sink.
Where is the data catalog in all this? Can a UX person survive in the world of events, objects, derived objects, data feeds ... and not go crazy (and who can help them)? Finally, why Sherlock wouldn’t have managed without Mrs. Hudson in the long run.
Watch the full talk
Watch this WaysConf session, then continue with related talks or explore the current programme.
Talk in brief
A case study in turning an organization’s scattered, poorly understood data into a usable product capability. The talk follows a team from the intimidating scale of raw events through interviews with the people who search, interpret, and govern them. It covers tribal knowledge, isolated retailer datasets, metadata, relationships between tables, and a discovery interface that makes ownership visible. The central lesson is that a data lake becomes valuable only when people can understand what exists, where it came from, how it connects, and whether they are allowed to use it.
Key takeaways
- 01
Make data scale understandable first
Physical comparisons turn abstract volumes into a shared problem that product, design, and business colleagues can discuss.
Watch from 2:26 - 02
Begin discovery with the people searching
Interviews reveal the workarounds, trusted colleagues, and decision risks hidden behind a request for a technical catalogue.
Watch from 8:20 - 03
Capture tribal knowledge where work happens
Search behavior across chat, documents, and personal contacts shows which context a formal data product must preserve.
Watch from 14:03 - 04
Respect commercial boundaries in the architecture
Retailer datasets require explicit separation and permissions so discovery does not expose competitively sensitive information.
Watch from 18:52 - 05
Treat metadata as a design responsibility
Clear event and table descriptions allow future teams to reuse data without reconstructing its meaning from the original authors.
Watch from 35:34
Video chapters
- 2:26The overwhelming scale of event data
The opening translates thousands of terabytes into concrete comparisons that establish the discovery challenge.
- 8:20Starting with people rather than infrastructure
Research focuses on the analysts and teams who locate, interpret, and depend on organizational data.
- 14:03Searching through chat and tribal knowledge
Existing workarounds expose where context lives and why a database alone cannot answer every question.
- 18:52Separating retailer data safely
The architecture must support discovery while enforcing commercial and regulatory boundaries between clients.
- 24:05Prototyping a shared discovery experience
A product concept shows how teams could browse ownership, meaning, and relationships in one place.
- 29:50Linking tables through visible keys
Identifiers and relationship cues help people understand how datasets can be combined responsibly.
- 35:34Designers as stewards of useful metadata
The conclusion asks product teams to document the data their own interfaces and analytics generate.
Read edited transcript highlights
These concise notes were edited from automatic captions and checked against the talk structure. They are not a verbatim transcript.
Scale needs a human frame before a technical answer
A number measured in thousands of terabytes is too abstract to guide a product conversation. Translating it into physical distance or years of continuous viewing helps colleagues recognize the scale without pretending the comparison solves it. The framing creates a shared starting point for deciding which parts of the data problem are most valuable to address.
Watch from 2:26The first users are the people navigating uncertainty
A data catalogue should begin with the analysts, product teams, and operators who currently search for information. Their work reveals which questions repeat, who is trusted, how quality is judged, and what happens when the wrong table is used. This evidence defines the product more accurately than beginning with a preferred database interface.
Watch from 8:20Informal channels contain essential context
People search chat histories, ask experienced colleagues, and rely on undocumented memory because a table name rarely explains purpose, ownership, freshness, or limitations. A formal discovery product should capture that context and make responsible contacts visible. It should reduce dependence on tribal knowledge without discarding the expertise behind it.
Watch from 14:03Discovery must not weaken data boundaries
A shared platform can serve several retailers while still keeping their commercial data strictly separate. Permissions, ownership, and lineage must be part of the experience rather than invisible backend rules. Users should understand what they may discover, what they may access, and why a restriction exists before they attempt an unsafe combination.
Watch from 18:52Product analytics also requires documentation
Designers and product teams create data through events, funnels, experiments, and interface behavior. Those records should be named and described so another team can interpret them later. Metadata is therefore not only a data-engineering concern; it is part of making the product’s decisions inspectable, reusable, and less dependent on the original team.
Watch from 35:34

