This blog post is AI-Assisted Content: Written by humans with a helping hand.
Documentation is for machines now. Understanding is for people later.
Albert Einstein famously said, “Never memorize something that you can look up in a book.” For years, we’ve used this as a high-minded excuse to justify our own procrastination. Why bother learning the syntax when (insert AI of choice) is right there? But the goalposts have moved. We don’t just have books anymore; we have infinitely more powerful tools like semantic layers and AI agents that can traverse a massive codebase and recall every edge case instantly. We are entering an era where, armed with these tools, merely not being a 10x developer can feel like procrastination.
This essay was inspired by that reality, and by several talks I heard at dbt Summit, especially Benn Stancil’s talk on how humans have ideas. He argued that human creativity relies on random collisions between disconnected memories, a leap AI’s neatly indexed models struggle to make. An example he presented was connecting a walk in Central Park and the many dogs present to the total absence of pets in Las Vegas. This was an idea he had when presented with the task of explaining Las Vegas to someone. Was this idea revolutionary? Perhaps not, but it didn’t need to be. As a result, I left with a bunch of ideas rattling around in my head. Rather than letting them disappear while walking to my nearby gas station for another Monster, I decided to write them down. What follows is my attempt to connect those ideas to the considerably less glamorous reality of data engineering: Thinking about problems, writing things down, organizing the resulting pile of notes and figuring out what AI means for all of it.
The Two Readers
There is a particular kind of data engineering problem everyone eventually encounters.
You’ve been staring at the same dbt model for an hour. There’s a join that shouldn’t be necessary, but exists because a source system was configured five years ago and nobody wants to touch it. A column that means three different things depending on which pipeline produced it. A dashboard that is technically correct, yet answers none of the questions anyone actually has.
You try changing the SQL. You try staring at the SQL. You try asking an AI to explain the SQL, which yields 600 words defining what a LEFT JOIN is.
So you stare at it some more.
It feels like work. It looks like work. But still it feels directionless because we misconstrue what the hard part actually is: Figuring out what problem the query was supposed to solve in the first place. When data teams attempt to solve this confusion, their instinct is often to dump every piece of technical metadata, pipeline logic and edge case into a massive wiki, effectively asking human engineers to process information like a compiler.
And this is why most data documentation fails.
For years, documentation has been the data industry’s ultimate unfulfilled moral obligation. Everyone agrees we need it. Nobody wants to write it. And when someone finally produces a 40-page guide, almost nobody reads it. We blame laziness or perhaps an inability or lack of desire to keep up with trends or the newest AI enhancement. But the real problem is that we keep pretending documentation has one reader.
It doesn’t anymore. It has two:
- A machine trying to enforce rules.
- A human trying to form a mental model.
The Machine Wants Everything
To an AI agent living inside an IDE, CI check or code review workflow, context is cheap and appetite is infinite. It does not get tired or build resentment when you give it 8,000 words of repository conventions, a table of naming rules, and it won’t open Slack halfway through a data contract spec and forget what it was doing.
For the machine, exhaustive documentation is good documentation.
If you want an agent to catch a broken lineage edge, reject an invalid model name, or notice that someone embedded business logic directly inside an exposed mart, you should give it the boring details. All of them. The schema rules, folder hierarchy, exception cases, ownership metadata and test expectations. Learning to communicate these technical specifications effectively so an agent can parse them is rapidly becoming a distinct engineering skill in itself.
An LLM does better when the rules are explicit, dense and unambiguous. In that sense, AI has made a certain kind of documentation much more useful than it used to be: The exhaustive kind nobody wanted to read.
Which is part of an ongoing problem. Despite how good we believe we are at processing information, humans are not machines.
The Human Wants a Map
When a human data engineer enters a complex data repository, they are usually not trying to ingest the full institutional memory of the company. They are trying to answer a smaller and more urgent question: “Where am I?”
A person does not need every edge case in the warehouse on day one. They need the shape of the system:
- What belongs in silver?
- What belongs in gold?
- Where does business logic live?
- What is the grain of this model?
- How do I not repeat a model?
- What kinds of models are exposed?
- What does “good” look like here?
- How do I know whether I am adding to the architecture or damaging it?
It is important to not misconstrue this as as needing less documentation. It is needing different documentation.
A human cannot brute-force their way to understanding by cramming a thousand disconnected truths into short-term memory. We learn by forming a mental model and an intuitive sense of the workspace. This is something practical to all areas of life, not just data.
A bad way to learn a city is to walk every street until geography eventually reveals itself. This works, technically. If you move somewhere new and spend enough time wandering, you will eventually learn that the river matters, downtown is east of the park, the train only runs north-south and the airport is much farther away than everyone implies. After a few weeks, you will know which roads connect, which neighborhoods are fake names invented by real estate agents and which intersections should be avoided if you value your time
But this is a strange thing to call onboarding and yet this is exactly how many engineers learn a data repository. They open a model, follow a ref to a macro, then a YAML file, then a model with the word “final” in its name that is naturally not final in any meaningful sense. The architecture was there the whole time, but the engineer had to discover it archaeologically. Or, worse, they attempted to force a pattern onto something that is actually just a chaotic mess, digging themselves deeper into confusion.
Good human documentation should provide the map before asking someone to memorize the streets. Without that map, mental models are fragile and do not survive well inside walls of text. This is why long documentation often produces the exact failure mode it was meant to prevent. An engineer reads three paragraphs, feels their soul leaving their body, closes the guide, copies a nearby model, and inadvertently violates three architectural assumptions they didn’t know existed, and goes back to fix them, at best saving no time, and worst, wasting far more. The documentation was technically correct, but useless.
Good Human Docs Help People Think
Technical documentation often overlaps with how engineers generate good ideas. Eventually, technical knowledge stops being the primary bottleneck, and the real challenge emerges: Knowing what to build.
As mentioned at the beginning, human brains excel at mashing wildly unrelated concepts together until something original shakes loose. This is a domain where we still outshine machines that merely optimize within closed rules. Because of this, my best ideas rarely come from staring harder at a screen — they come from unexpected collisions. Good human documentation should exploit this: It shouldn’t merely specify a system, but give your brain something strange enough to push against so you can think with it.
You are thinking about centralized data architecture, and then you watch Carmy try to run the pass in “The Bear.” You are thinking about semantic layers, and then you read about how Tokyo subway stations use distinct musical chimes instead of loud alarms to prevent crowd panic. Will everyone draw the same conclusions on how each pair relates? Absolutely not, but it is also not necessary. It is in this area where we can better understand how we build our mental maps and understanding of any topic for that matter, and when iterated and combined with other’s ideas, we start to develop something truly robust. Their job isn’t to be perfect models; their job is to trigger insight
Translating a sprawling technical specification into a map that sparks these kinds of connections is a distinct engineering discipline. Anyone can transcribe a database schema, but structuring context so that people can reason through complex problems requires real craft. We tend to be overconfident in this skill. I certainly can be at times. Because of this it is important that we remind ourselves that it is work that shouldn’t be rushed, and it deserves to be revisited often.
The Small Set of Rules
This is why I like documentation that looks less like an encyclopedia and more like a set of constraints. Silja Märdla’s dbt Summit talk about Bolt’s explicit data modeling framework is a great example. If you have 15 minutes to spare this week, it is genuinely worth your time to take a look. What I enjoyed was that Bolt didn’t tackle their issue by simply writing more documentation as many of us have done in the past, but that they reduced the system to a small number of structural rules humans could actually carry around in their heads. They made it simple and memorable.
- Exposed models have specific shapes.
- Transformation logic lives in specific places.
- Intermediate models play recognizable roles.
- Metrics are separated based on how they can safely aggregate.
This kind of documentation is useful because it gives people an architectural compass. It does not explain every possible implementation detail. It explains how to think. That distinction is crucial, as a machine wants the full spec because it is trying to enforce behavior whereas a human wants the map because they are trying to make judgment calls.
- Should this model exist?
- Should this be a dimension or a fact?
- Is this logic reusable?
- Are we solving a data problem, or are we modeling around an organizational misunderstanding?
And So…
The future of data documentation is not “better docs” by improving our prompts or bringing more people in. It is split docs.
For the machine, we need exhaustive, generated, formal, code-adjacent documentation. The machine gets the manifest, schemas, contracts, lineage graphs, policy rules and edge cases. For the human, we need concise, visual, architectural documentation that creates understanding — giving them the shape of the system. (The machine docs can always serve as a deep-dive appendix for humans later, but they are never the starting point).
This sounds obvious until you notice how often we do the opposite. We ask humans to read machine documentation, then act surprised when they skim it. We ask machines to interpret human prose, then act surprised when they hallucinate around ambiguity. A single document called “Data Modeling Standards” expected to both onboard a new engineer and instruct an AI code reviewer is doomed.
Doc Rot Gets Worse With AI
If writing documentation is the tech industry’s favorite unfulfilled moral obligation, maintaining it is our most spectacular collective delusion. And before we pat ourselves on the back for elegantly solving the crisis with this new two-reader paradigm, there is a rather significant catch: Machine documentation rots faster and with infinitely more collateral damage than human documentation.
If a human reads an outdated wiki, they usually figure it out the moment their first query fails. If an AI agent reads that same outdated page, it will undoubtedly synthesize 400 lines of impeccably formatted, entirely invalid SQL, hallucinate three macros, and open a pull request. Letting an AI rely on verbose, “stateful” documentation like hardcoded Jira references or manually snapshotted project structures in a context file is essentially laying landmines in your own repository. When the implementation inevitably drifts, the AI will gaslight itself into trusting its stale notes over the actual source of truth.
The takeaway here is that machine documentation can no longer depend on human virtue for updates, which is a maintenance strategy that has historically had a zero percent success rate anyway over a long enough period of time. People get their priorities shifted, things get forgotten and people leave the company, taking their context right along with them. Machine context needs to be derived, not drafted or memorized. Schemas must be pulled directly from dbt artifacts, contracts and CI outputs, while tests pull double duty as executable documentation. If you want to actually help the agent, your context files should guard the structural “why.” By this I mean the architectural constraints and non-negotiable design patterns while refusing to snapshot the “what.” We are crossing into an era of workflows where the most reliable machine documentation isn’t written at all. It is compiled.
Human Docs Should Rot Differently
Human documentation has a completely different, and arguably more deceptive, failure mode. It rots when the code changes, sure. But it also rots the second it stops being a useful mental model.
This creates a false sense of security, because a human guide can be technically accurate while still being fundamentally broken. If reading it doesn’t give someone an intuitive sense of where a pipeline belongs or why the architecture looks the way it does, it has failed. We need to stop treating human documentation as a junk drawer and start maintaining it like a product interface. Can a new engineer figure out the blast radius of the system in ten minutes? Can they intuit where a change goes?
The ultimate litmus test is whether they can ask better questions after reading it. The best human documentation wont just spoon-feed facts that have to be referenced over and over again, but rather it equips people with high-quality questions to carry around. “Why does this table exist at all? Why does every department have a conflicting definition of ‘customer’? How does this asset help solve a business objective?”
The New Contract
We are entering a world where AI makes technically competent answers aggressively cheap. Mostly, this is a good thing. But when answers cost nothing, knowing which questions to ask becomes the only skill of value.
A machine can flawlessly optimize a data pipeline, but a human still needs to wonder if that pipeline deserves to exist in the first place. A machine will blindly enforce a data contract to the letter, but a person must decide if that contract is quietly lying to the business. A machine can memorize every rule in the repository, but a person needs a map concise enough to realize when those rules are pointing at the completely wrong problem.
Fantastic! But a compass is useless if true north keeps shifting, and a rulebook cannot govern fundamental chaos. Like many others, I often hope that reading an article somewhere will grant me the secret sauce, but this two-reader paradigm is not a quick fix you can bolt onto a broken warehouse. It only works if you are making a forward-looking investment in the foundations that actually define your business logic namely, clean semantic modeling and resilient data governance with collaboration and agreement across teams. Without those, an AI agent will simply help you enforce your mess at scale. (If you’re trying to figure out how to lay that structural groundwork instead of just applying another band-aid, that kind of architectural engineering is exactly what we do at InterWorks).
We are operating under a new contract where our machines get the entire detailed rulebook, and we just use the index. The best engineering cultures will be the ones that finally stop confusing the two.
