CORE DIRECTORY // SYSTEM.USER.DIANA_ISMAIL
Labs by Diana — Experiments that ship.
Side projects that got out of hand. AI tools built for problems I kept tripping over — now live, now yours.
The Paradox AI Made Worse
ARTICLE_063
PUBLISHED
2026.09.09
READ
~11 MIN
Polanyi's paradox is usually invoked as an automation story — the tasks a model still can't do because the person doing them can't write down the rule they're following. That's true, but it undersells the problem. AI hasn't just left tacit knowledge untouched at the edges it can't reach. Inside the territory it has taken over, it has made the paradox worse: when generation was expensive, the process of producing something was slow enough to leave traces of the reasoning behind it almost by accident. When generation is cheap, the artefact arrives finished, and the judgment that shaped it is gone the moment the output appears. Two structural answers exist, from opposite directions. One captures the reasoning after the work — a record of situation, decision, risk, and change, written down once the call has been made. The other captures it during the work — running the work somewhere a second person can watch it happen, before it's finished, while the judgment is still visible. Neither replaces the other. Most organisations, including most AI-heavy ones, currently have neither.
This piece works through Polanyi's actual claim, why cheap generation sharpens rather than dissolves it, what each of the two capture strategies buys and where each one runs out, and what it looks like to run both at once — not at Shopify's six-thousand-person scale, but at the scale of one operator and a working set of files.
The_Machinist_Who_Can't_Say_What_She_Knows
As Nate tells it — and I'm relaying his account rather than something I've verified independently — there's a machinist in Oregon responsible for testing a specific type of screw used in a Boeing assembly. She can tell, by feel, whether a given screw will hold under the conditions it needs to hold under. Nobody, including her, has managed to write down exactly what she's checking for. When she retires, the company doesn't lose someone who follows a documented procedure. It loses the procedure itself, because the procedure was never separable from her hands.
Whether the details are exactly right matters less than what the story is doing as an example: it's the sharpest, most physical version of a problem that shows up anywhere expertise accumulates through practice rather than instruction. The knowledge is real. It's reliable. And there is no clean way to get it out of her and into a document, a training programme, or a model.
What_Polanyi_Actually_Said
Michael Polanyi named this directly, in 1966: "I shall reconsider human knowledge by starting from the fact that we can know more than we can tell" (The Tacit Dimension, p.4). His own examples were smaller than a Boeing screw — we recognise a face we know well without being able to describe the arrangement of features that let us recognise it; we ride a bicycle without being able to state the physics of staying upright. The knowledge is real, it's reliable, and it doesn't pass cleanly through language.
Decades later, the economist David Autor picked the same paradox up as an explanation for what automation could and couldn't reach: abstract tasks needing creativity and open-ended problem-solving, and manual tasks needing situational adaptability, both resisted automation for the same reason — nobody could specify the rule a machine would need to follow, because the person doing the task had never had a rule to specify in the first place. Autor's version is where most people meet Polanyi's paradox now: as an explanation for occupational polarisation, why the middle of the skill distribution emptied out while both the top and the bottom held.
Why_Cheap_Generation_Makes_It_Worse
The automation version of the paradox treats it as a boundary — a line marking what machines can and can't do, moving slowly as capability improves. That framing misses something that's become obvious from inside an AI-heavy workflow: the paradox doesn't stay outside the boundary. It gets worse on the side machines have already taken.
Before generation was cheap, producing something — a report, a piece of code, a plan — took enough time and friction that some of the reasoning behind it tended to survive in the process: the draft you could see the previous draft underneath, the comment explaining why an approach was rejected, the conversation that happened before the document existed. None of that was designed as a record of tacit judgment. It was a side effect of generation being slow. Generation is fast enough now that a finished, polished artefact can appear with none of that trail attached. The model doesn't leave a rejected-approaches comment. It doesn't need an afternoon of thinking-out-loud on the way to the final version. It produces the final version directly, and the reasoning that a slower process would have left visible almost by accident never gets externalised at all, because nothing forced it to.
This is the sharper version of Polanyi's problem. It isn't that AI can't do the machinist's job — plenty of AI commentary stays fixated on that boundary. It's that AI has made the artefact a worse witness to the judgment behind it, across the entire territory it now touches, including the territory it's genuinely good at. Fast, cheap, competent output is precisely the condition under which tacit knowledge stops leaving traces.
Two_Ways_to_Capture_What_Can't_Be_Told
Two structural answers exist, and they capture the same kind of thing from opposite directions.
The first is post-hoc: write the reasoning down after the judgment has been made, in a fixed structure, so a reader who wasn't there can reconstruct why the call was the right one. Nate's talent board proposes exactly this — a record of situation, decision, risk, and what changed as a result — as a response to a specific failure: when AI made output cheap, a finished piece of work stopped being evidence of the judgment that produced it, so the only credible evidence left is a deliberately captured account of the reasoning itself. The strength of this approach is that it's cheap to build and durable once it exists — a record survives the session, the person, the model that produced it. The limit is that it's only as honest as whoever fills it in after the fact, and it captures the conclusion of a judgment more reliably than it captures the judgment while it was still in motion. You get the situation, the decision, the risk, and the change. You don't get the moment the decision could have gone the other way.
The second is in-flight: don't wait for the work to finish — put the work itself somewhere a second person can watch it happen. This is what Shopify's River does by design: the agent runs in public channels rather than private messages, so a colleague can watch a senior practitioner frame an ambiguous problem, push back on a suggestion, and correct a wrong turn, in real time, before the finished output exists to obscure any of it. The strength is that it captures judgment nothing else can — the moment of correction, the reasoning behind a rejected approach, the version of the work that never shipped. The limit is that it only works if someone is actually watching, and it doesn't produce a portable record the way a written log does. If nobody was in the channel when the correction happened, the moment is gone just as completely as it would have been in a DM.
The_Architecture_That_Uses_Both
I don't run a talent board and I don't run River, but I run structural equivalents of both, and I hadn't connected them to each other — or to Polanyi — until writing this.
decisions.md is the post-hoc version. Every entry follows roughly the talent board's own shape without having been built to match it: a decision, the reasoning behind it, the alternatives that were weighed and rejected, and what actually got implemented. It's written after the call has been made, by whoever made it, which means it carries the same limit the talent board carries — it's a record of the conclusion, not a transcript of the deliberation. What it buys is durability. A decision made months ago is still legible today, to a session that had no part in making it, in a way a private conversation about that decision never would have been.
The Mano persona system is closer to the in-flight version, at a much smaller scale than a public Slack channel. Every session starts as a base layer with no role-specific guardrails, and when the work calls for a different mode — dispatching several specialists, or drafting in a particular voice — the shift is announced inside the session itself: [ENTERING SHEENA MODE], [ENTERING REID MODE]. That's a small thing, but it's the same structural move as River's no-DM rule in miniature: instead of the judgment about which mode fits this moment happening silently and being reconstructable only from the output, it's stated at the point it's made, in a form anyone reading the session afterwards can see. It doesn't capture everything — most of the reasoning inside a persona shift is still not externalised — but it captures the fact of the shift and the moment it happened, which a purely post-hoc record never would.
Neither one is complete on its own. decisions.md would eventually accumulate a set of conclusions with the reasoning between them thinning out over time, if nothing ever showed the moment-to-moment judgment behind a session. The in-flight announcements would eventually be noise nobody could search, if nothing ever consolidated them into something durable and readable months later. Running both is what makes either one worth having.
What_Gets_Lost_Otherwise
Polanyi's line is usually read as a limit on what can be automated. It's better read as a limit on what any system — human or otherwise — can transfer without deliberately building a way to transfer it. The machinist's knowledge didn't disappear because nobody valued it. It disappeared because nothing was built to carry it past the room she was standing in. AI didn't create that condition. It just made it cheaper to produce work that looks finished without ever having been asked to explain itself, which means the absence of a transfer mechanism now costs more, faster, than it used to. Building one — after the fact, in flight, or both — isn't optional infrastructure. It's the only way anything survives past the moment it was known.
KEY_TAKEAWAYS
TAKEAWAY_01
Polanyi's paradox — we know more than we can tell — isn't a boundary marking what AI still can't do. Cheap generation has made it worse inside the territory AI already covers: a fast, finished, competent artefact carries none of the trail a slower process used to leave almost by accident, so the judgment behind it disappears the moment the output appears.
TAKEAWAY_02
Two structural answers exist, from opposite directions, and each has a real limit the other one covers. Post-hoc capture — a record of situation, decision, risk, and change, written after the call — is durable but only as honest as what someone chose to write down afterwards. In-flight capture — doing the work somewhere a second person can watch it happen — captures the judgment itself, but only for whoever happened to be watching, and produces nothing portable once the moment passes.
TAKEAWAY_03
Running both is not redundant. A post-hoc decision record and in-flight, in-session announcements cover for each other's specific failure mode — one thins into conclusions with no visible reasoning if nothing shows the moment-to-moment judgment, the other becomes unsearchable noise if nothing consolidates it. An organisation, or a one-person fleet, needs a structural answer working in both directions, not a choice between them.