Documentation, Disrupted: How Two Technical Writers Changed Google Engineering Culture Google

Speaker: Ríona MacNamara, Staff Technical Writer, Google Event: Write the Docs


1. Introduction: The Winding Road to Tech Writing

Hi, I’m Ríona MacNamara. I’m a staff writer at Google—I’ve been there about eight years. Last year was my first time at Write the Docs, and what a year it has been! It has genuinely been the best year of my career, which spans about 20 years in technical writing. In terms of what we’ve accomplished and the team we built around internal engineering documentation at Google, it all started right here.

I love conferences like this because technical writers come from such diverse backgrounds. My own path to tech was definitely not direct:

  • I started off studying English because I wanted to write the Great Irish Novel. By graduation, it became clear that wasn’t going to happen.

  • Next, I went to law school thinking I’d be the "people's hero." By the end of law school, that clearly wasn't happening either.

  • Then I moved to London to join literary publishing. I imagined nursing great novels into existence and attending parties with Salman Rushdie. Instead, I landed at Virgin Publishing, where I worked on books about Doctor Who, erotic fiction, and serial killers. (I’ve forgotten more about serial killers than you will ever know!)

  • Later, I worked at Random House and then moved to Dublin to be managing editor at Attic Press, an influential feminist publishing house.

None of this had anything to do with technology until I took a two-week freelancing contract at Microsoft in Dublin editing content for Encarta World Atlas. That two-week gig turned into 10 years at Microsoft, where I eventually became International Editor for Encarta and moved to Redmond in 1998. After Microsoft, I spent 18 months at Amazon, and now I’ve been at Google for eight years.

I love my job. I’m one of those intolerable morning people who bounces out of bed before coffee, excited to open my laptop.

Except, this time last year, that wasn’t the case at all.


2. Burnout and the Science of Happiness at Work

A year ago, I was feeling deeply burned out. I kept wondering: Am I making an impact? Does my work actually matter?

Around that time, I came across Shawn Achor’s TED Talk, “The Happy Secret to Better Work.” We are raised to believe that if we work hard, success and happiness will follow. But Achor explains the reverse is true: happiness comes first. When we are happy:

  • We are more creative, connected, and productive.

  • Doctors make accurate diagnoses 31% faster when in a positive state of mind.

Achor outlines three key predictors of workplace happiness:

  1. Optimism: The belief that your work matters, makes an impact, and is recognized.

  2. Perception of Stress: Viewing stress as an energizing challenge rather than a paralyzing threat.

  3. Connection: Feeling deeply connected to your colleagues, your work, and having influence over direction.

At the time, I was failing on all three counts.


3. The Documentation Crisis at Google

To understand why, you need to understand the scale of Google’s internal documentation challenge:

  • 23,000 engineers.

  • Hardly any internal technical writers. The ones we have are brilliant, but the reality is that engineers have to write and maintain their own documentation across thousands of shifting, interdependent systems.

  • A fiercely autonomous engineering culture: Nobody can tell a Google engineer what to do. They choose their own tools, their own systems, and they hold very strong opinions.

The result was total fragmentation:

  • Documentation lived scattered across Google Sites, Google Docs, wikis, and random HTML pages.

  • Much of it was unmaintained, unreliable, or nonexistent.

  • In Google’s annual company-wide survey—which leadership takes very seriously—the #1 blocker to engineering productivity two years in a row was the lack of discoverable, reliable internal documentation.

We had tried fixing this before:

  • Top-down mandates didn’t work (engineers ignored them).

  • Bottom-up efforts didn’t scale (we’d polish a team’s docs into an exemplar, but the moment we walked away, the docs decayed).

On top of that, I had recently been promoted. Instead of feeling relaxed, my imposter syndrome spiked. I felt isolated, stressed, and convinced my efforts weren't moving the needle.


4. The Write the Docs Epiphany

Then, I came to Write the Docs last year.

Two talks in particular stood out, especially one from Twitter explaining how they co-located documentation with code, alongside another talk on Minimum Viable Documentation.

I remember furiously instant-messaging my coworker, Aaron, in our New York office:

"They have the exact same problem we do, and they solved it by putting docs in the repository with the code!"

When we got back, Aaron and I—both between projects—decided to dig into the problem from scratch. We interviewed as many engineers as possible. We quickly realized our core mistake:

All our previous documentation solutions had been designed for writers. But technical writers weren't the primary users—engineers were.

The root issue was architectural:

  • Google has one of the largest integrated monorepos in the world, with tens of thousands of engineers committing to it every day.

  • Code lived in this tightly managed, unified repository.

  • Docs lived everywhere else: in a Google Doc, on a wiki, or on a sticky note.

Documentation would never become part of Google’s engineering culture until it lived in the codebase and integrated directly into the engineering workflow.


5. Going "Five Blades": The Birth of g3doc

We didn’t want to build just another competing content management system. Inspired by The Onion’s satire piece ("Fuck Everything, We're Doing Five Blades"), our code name became Five Blades: we were going to go all-in or go home. We wanted a system that could become the standard for all 23,000 engineers.

Our vision was radically simple:

  1. Allow engineers to write docs in Markdown right alongside their code in the codebase.

  2. Render those files automatically at a clean, predictable intranet URL.

  3. Keep the barrier to entry near zero.

There was one technical catch: Google’s internal doc-serving system couldn't read directly from the core codebase.

Having nothing to lose, we emailed Steve, an engineer in Munich who ran that infrastructure: "Hey, could we render docs straight out of the repo?"

Because of the time difference, I woke up the next morning to find that Steve had already written a design doc, scheduled a security review, and laid out the architecture. Within nine days, we had a working proof of concept.

Project g3doc was born.

Our branding was intentionally minimal: just a terminal prompt. Drop Markdown files into your directory, add a simple sitemap.md for navigation, and g3doc rendered code formatting, dynamic tables of contents, and page metadata out of the box.

On July 8th, we launched our first live site for a core infrastructure system. The developer reaction was immediate:

"This is the best thing ever."


6. Viral Adoption and Culture Shift

Within weeks, 18 projects migrated. We only actively pitched three of them; the rest spread purely by word-of-mouth.

Other technical writers facing massive scale joined the effort—Ed from Search Infra, Ricardo from Ads, and Theodore, who was supporting a team of 300 developers across dozens of projects. We operated with open-source principles: governance belongs to the contributors.

We established two core cultural rules for our team:

  1. Focus relentlessly on the engineer. Every feature had to respect developer workflows.

  2. No "cookie licking." (Cookie licking is when someone claims an interesting feature by touching it, but never finishes it, preventing anyone else from doing it.) We prioritized momentum, rapid iteration, and immediate impact.

Next, we partnered with Google’s Developer Infrastructure team to embed g3doc into the core developer toolchain:

  • Code search

  • Code editor

  • Code review tools

Engineers could now preview rendered documentation directly within code reviews with a single click.

The Numbers:

  • September: 18 projects

  • December: ~160 projects

  • Q1 Goal: 300 projects (an aggressive OKR)

  • Actual Q1 Result: 930 projects

  • Today: Over 1,600 projects

Most of the team building and scaling g3doc were 20% volunteers. It evolved from a scrappy grassroots effort into core company infrastructure.


7. Lessons Learned

1. Focus on the engineer’s workflow, not the writer’s.

In an environment where engineers own the docs, you cannot expect them to adopt a technical writer’s toolset. Meet them where they already live: inside their repository, editor, and review tools.

2. Momentum matters: Start small and iterate.

Don’t over-engineer a massive CMS upfront. Build a rock-solid, minimalist base platform, launch it quickly, and add features based on real developer demand.

3. Claim authority through solving real problems.

Between us, our team had decades of technical writing expertise in information architecture, layout, and style. But none of that mattered until we solved the engineers' core problem. Nobody was going to invite us to "fix their docs."

Once we fixed the workflow, we didn’t just gain adoption—we gained influence. Senior engineering directors began reaching out, funding our efforts, and inviting us to strategic architectural discussions.


8. Closing: A Happiness Check-in

Revisiting Shawn Achor’s predictors of happiness:

  • Optimism: Massive. Thousands of changelists (CLs) are submitted every day where engineers update code and documentation in the exact same commit. Docs are finally part of the engineering culture.

  • Perception of Stress: I work hard, but I feel less stressed than ever because the work is self-directed, purposeful, and demonstrably impactful.

  • Connection: Unprecedented. We moved from isolated writers to trusted partners seated at the table with principal engineers.

A few months ago, Aaron and I were in a meeting with senior engineering leaders in New York discussing a major codebase-wide integration for g3doc. One engineer noted, "This might require a global policy shift across the monorepo."

A principal engineer leaned back and said:

"Well, three of the five people in this room are global approvers of the codebase. I don't think that's going to be a problem."

Right then, my phone buzzed. It was a text from Aaron sitting across the table:

"We did it."

Authority and influence are out there for technical writers—they are ours for the taking if we focus on our users, step up, and solve real engineering problems.

Go five blades. Be happy. Thank you!


Key Takeaways Summary

  • Docs as Code Works: Placing documentation in the same repository as the code ensures docs are updated, reviewed, and versioned simultaneously with code changes.

  • User-Centric Tooling: For internal engineering docs, the engineer is the user. Markdown + Git/Monorepo + existing review tools will beat an external wiki or CMS every time.

  • Grassroots Credibility: Don't wait for top-down mandates. Build a lightweight, reliable prototype, prove value, and let peer adoption drive culture change.


n5321 | 2026年10月7日 20:04

TDD, Where Did It All Go Wrong

Um, just a quick sort of show of hands, or a couple of questions from the audience: How many of you have actually seen a video of me doing this talk at any point? A few of you, okay. There are some nuances, and giving it live will be a bit different, but you're seeing the gist of the argument. 

How many of you are using Test-Driven Development today? Test-first or test-after? Okay. How many of you are doing it test-first? All right. And how many of you are using mocking techniques and mocks? Okay, loads of you. Okay, very interesting.

Okay, this is me. Most of that is dull and uninteresting, but what I usually draw attention to is the last point on that slide: people like me come and talk at conferences to stand in front of you—we're not necessarily smarter or cleverer than any of you guys. And you know, I'd really encourage all of you to get up and speak at conferences if you can, contribute back to the community. That's all what I'm doing. I'm just a developer like you who decided to try and share some of my ideas with everybody else.

I kind of got forced into it because back when .NET first started, there were very few people who were the experts. And so, as such, you know, if you wanted to run a meetup or a user group, you had to kind of start up and do it yourself. But I'd encourage all of you to get involved if you can. It's a great way of personal growth for you as a developer in terms of learning to express your ideas more clearly. 

That's who I work for; I work for Huddle. It's not really that exciting—we are hiring, by the way—but it's not that really exciting for you other than just to say we are a SaaS software business. I have nothing to sell you, right? You can't buy my services. So, all of the ideas here are based simply on my experience I'm trying to share with you, not on me trying to basically market anything to you.

What we're going to talk about today:

I want to talk about what I perceive as being the problem with how people are practicing TDD today and why there has been growing resistance to TDD. 

I want to then look at rebooting our TDD practice, taking us back really to where we started from, and looking again at ideas like Red-Green-Refactor, clean code, refactoring, etc., to understand how those ideas should be applied, where we may have made mistakes in our practice that have led us to adverse ways of doing TDD. 

And then I'll talk about some ideas that can help: Ports and Adapters, gears, look at ATDD, BDD, and mocks. We want to get through some of the later sections; we may go a little bit faster depending on time, so we'll try and focus on the early sections.


First thing I'll do is talk about the problem. 

Now, I have been practicing TDD since about 2004 in London. The .NET community were quite early adopters of Test-Driven Development, just simply because we had close relationships with an XP group that used to meet in London and a lot of cross-fertilization of ideas. And so we were into practicing TDD for quite a while. 


And I first gave this talk in about 2013, so I'd been roughly 10 years doing TDD. And I had come to the conclusion, after owning TDD-produced test suites for a while, that there were some problems with what we were doing. And you know, one of my things is that I tend to try, because as you become a software architect, to stay at companies long enough to see my own mistakes come back to meet me. It's a very humbling thing to do, but it's also how you gain kind of a lot of value from your experience. If you stay somewhere two years, quite often the problem is that you never see... you finish implementing something, you're very proud of this great work you've done, and you move off. And you never find out what were all the things that you actually did wrong that later developers whined about. So if you stay a bit longer, quite often you find what goes wrong. And I had written quite extensive test suites, and I found some of them were very difficult and very expensive to own. So that's going to talk about my understanding of TDD and how that shifted over the years.

Okay, so the first thing was back in 2005–2006, when I had first drunk the Kool-Aid on TDD. When I was practicing TDD, I found quite a lot of resistance to Test-Driven Development. A lot of people were saying, "We don't want to do this. This is a crazy idea." And I was very much, "You just don't get it. You just don't get what the rest of us get. You know, this is the best change in development practice I've seen in the last 10 years. You guys need to be doing TDD. You just don't get it." 


But still, the people saying, "Well, I just... I'm not really sure about this TDD, this crazy thing." A lot of those guys were really quite smart developers who I respected. You know, these are guys who wrote really good code. Yet they were saying to me, "This TDD thing, it's not the right way to develop software." And you know, the more they said this, the more I just kept repeating the TDD mantras, as if like a child—if I just kept speaking slowly and carefully and repeating how TDD worked, eventually the light bulb would go off and they would appreciate the glory that was TDD. 


But a few of them just seemed to resist. I remember one of them, a really smart developer who I really, really admired his work, said to me one time, "Okay, and I finally got it." And I was like, "Yes?" And he said, "The junior developers—they need TDD. It really helps them. I can see how it helps them write good code. But surely you don't really expect the experienced developers on the team to be doing TDD, right?" And I was really puzzled. I couldn't really understand, what is this guy getting at?


A couple of times, we were very proud of the fact that we produced some code that went into production with almost zero defects. There was no UI involved, making it a bit easier, but one time I worked on a backend service and that went to production, and for a year there were no reported defects. Worked exactly as expected. We were very proud, and we put it down that this had worked with TDD: we had a fantastic test suite, we knew all the scenarios we were going to operate under, and provably our code worked.


All I looked at this test suite, one of the things I noticed was that we had about three times as much test code as production code. And I can remember being a little bit, "Mmm, is this right?" So I went onto Twitter, source of all answers to all questions in life and universe, the source of truth and wisdom. And I said to Twitter, "Twitter, is this right?" "Yeah, no, no, we quite often have, you know, two or three times the amount of test code that we have basically as actual production code." I was thinking, "Okay, but it feels a bit puzzling."


One of the problems with this is you quite often got the situation where, you know, some project manager would turn around and say, "Guys, could we just skip the tests? We'd be done so much faster. We'd get out to market a lot faster. I know you guys like doing good engineering, etc., I really get that. But it's taking us an awfully long time to write all these tests. Couldn't we just ship the code? It wouldn't really matter if there were a few defects, you know, in production, right? Can we just not do the tests?" 


And we'd say, "No, no, no, you don't understand. The way we write code is by writing tests. We can no more stop writing the tests than we can..." Let me collapse this all right into the standard mantra: "Tests are part of our practice for writing code. You're just telling us to stop writing code if you do that." And they'd go away and go, "Okay, fair enough, we get it, you know, you guys want to write tests." 


But we were slower. I mean, in reality, we were, you know, 20% to 50% slower than we had been before by writing tests.


The other thing I noticed over time was with so many big test suites is that when we wanted to change something, we'd quite often break a lot of the tests. And quite often what we were doing was simply refactoring. We were going, "Oh, well, we've added some new features to the code, and now the shape of the implementation is a bit wrong given these new features. And I can think of a better way of organizing our code to make it more maintainable. So I'm going to refactor the way the code is organized. I don't want to change the behavior of the system; I just want to refactor the shape of the code so that essentially it's more efficient." 


Wham! All the tests would break! And all the time, the tests that broke were the ones that basically were heavily mocked. Because these mocks were telling us exactly what ought to be happening in the code under test. They were saying, "Oh, first you should be calling this twice, then call this, and then call this and put these parameters on it." When we refactored something, we quite often changed how that worked. All these tests with mocks, they broke, and we had to rewrite all the mocks.


It was kind of puzzling, right? Because the promise was that we should be able to refactor our code, keep it basically fit and clean, we wouldn't get into a big ball of mud craziness, we wouldn't really see our code rot before us, because our tests would enable us to refactor, and refactoring kept our code healthy. But when we wanted to change our code, actually the tests were an obstacle to that change because they increased the amount of effort we had to put into making any changes. We had to fix all these tests that were breaking when we changed our implementation details. And that didn't sound right, because surely what Martin Fowler and Kent Beck were telling us about refactoring was that refactoring was changing the implementation details without breaking tests! Yet our tests were breaking all the time.


The other thing we began to notice was, I guess it's about 2009–2010, a whole new methodologies began to emerge that explicitly said: "Don't do Test-Driven Development, because Test-Driven Development slows you down." 


So there's Programmer Anarchy. You guys know what Programmer Anarchy is? Uncle Fred George—Fred George essentially said (for those who understand *Spinal Tap* references), he said, "If XP is good, let's turn the dials on XP up to 11. Let's do Extreme Extreme Programming. Let's basically fire everybody who isn't a software developer: fire QA, fire the product guy, just have developers talking to customers. And what we'll do is we'll write really small pieces of code, we'll call them microservices, they'll only be about a hundred lines." So different schools of microservices—this is the original one: about a hundred lines of code. "And we won't have any tests because it's such a small amount of code you're bound to get it right. We'll put it into production; if you want to change, we'll just throw it away and rewrite it, okay?" That's what's called Programmer Anarchy.


Then there were other things like "spike and stabilize," even Lean Software Development, which became the real, you know, go-to approach of California and Silicon Valley startups, were saying: "Don't do tests, right? Work directly with customers, get it out there fast as you can, get feedback, because you may be building the wrong thing altogether and going to throw it away. TDD is a kind of practice you do later on when you're scaling. Up front, tests slow you down too much." 


So there's this whole pressure saying TDD is a bad idea because TDD really slows you down and you're better off just getting feedback.


You guys especially over here: Duct-Tape Programmer. So the duct-tape programmer was this really annoying guy, right? And every organization has one. And he's the guy that business or the product team actually love, because he just gives them solutions as fast as he can. And as a requirement comes out, they say, "Let's give it to Bob. Bob will just get it out the door really fast." And the rest of the team are going, "But this is crazy!" The engineering team going, "We hate Bob! Bob's engineering is rubbish! We spend all this time maintaining the stuff that Bob's written and having to rewrite it!" And it's really annoying because Bob gets all the credit, right? From the product guy, the guys going, "Bob's fantastic! He's so receptive to our ideas, he pushes stuff out!" The rest of you are going, "But it's so hard to maintain, Bob's code is terrible, right?" 


But they love Bob, because Bob delivers for them. You're still writing your test suite, and Bob has shipped the code and it's running, okay? We had a few of those, and it used to be really annoying. It used to be really annoying that we were still doing our best engineering practices, and these guys were shipping code—and it was awful code—but they were shipping, right? And so they won the credibility, not us.


Sometimes we'd come back to our tests when they were breaking, quite often because we were trying to change implementation details, we'd realize that we had no idea what the test was doing. Anyone have this experience? You go back and find the test is red, and you're looking at it and you're going, "I have no idea what this test is for, and I have no idea whether the fact that it's red is good or bad at this point in time, okay?" This test seems to add two numbers and come to a result which I can't possibly understand how they could have originally ever got that result, and why this test was originally passing, and now I've refactored it and it's red—surely just because the test is wrong? 


Lots of tests just don't really give you any information, especially the ones that are like swarming with mocks, and you're kind of like, "What is actually under test here?" Anyone have the experience: I'd look at the test and then finally realize nothing is actually being tested apart from the mocks, right? Someone's written this thing around it so heavily, you go, "There's no real code actually under test here, okay?" And that experience basically began to drive us mad. We're like, "Well, I no longer know what this test is for. The best thing for me to do is just to delete this test because it's broken, I don't know why."


Made to my shame, a team I worked for embraced FitNesse, and we sat down and produced basically these living tables in HTML, and they drove... used them to drive our software. And eventually, you know, these little tests would show up, the page would show up green, so that we could say, "Yes, these high-level tests that are written in natural English so the customer can understand them, these tests basically demonstrate that our software does exactly what we say it does on the tin." And SpecFlow, Cucumber—these tools are all the same kind of thing: the idea essentially that I will write tests that the customer can just flip through the HTML pages, see the requirement we've implemented, feel confident that the system does exactly what he specified, because he can press the button and see it all go green. "Ah, the software runs exactly as I expected."


The developers basically hated writing these tests because they'd start out practically... they'd be red in the beginning of the iteration, so they'd be red for a lot of the iteration until eventually they would go green when you'd done the implementation. And that meant essentially your FitNesse test suite was quite often—well, correctly—red much of the iteration, so you could never really be sure if the tests were red, is that because it's yet to be implemented or is that because essentially we've broken it? And so the thing that happened was a lot of time the tests were red because we just assumed, "Oh, they're red because not implemented yet," when in fact they were broken. So they spent a lot of their time as broken test suites.


The customers never really looked at this stuff, right? We'd built these customer-facing tests, but it was very rare to get the customers to actually come along and flip through the HTML pages and say, "I'm so impressed, guys! I can see exactly the requirements that I want, I can see them written in a natural language, and I can see that it works, right?" Occasionally we'd try and sit them down, make them look at the screen, and they'd go, "Maybe send me a calendar invite, guys, and we'll talk about this at some point in the future, right?" And they were completely uninterested in seeing these test suites that we'd written to prove to them that the code works correctly.


And the check-in gates became more complicated, because as well as passing developer tests, you had to run all these FitNesse test suites before we allowed you to check your tests into the code base. They took a while to run; they were slow. And that meant essentially you'd go for making a cup of coffee and come back, essentially because, really, you know, that's you... then you'd lose track of where you were, what you were doing. And the whole flow of your coding session was broken, because now you're in that classic developer situation where it takes you 20 minutes to really get your head around the code that you're working on, and then you're in the groove and you're working on it, but now a forced interruption to our process is going to mean that you stop working effectively during that time period, okay?


So it became a really inefficient way of doing it, and the developers just really hated doing it and they started to ignore them, right? "If we just basically keep going and checking code in, eventually someone will notice the tests are all red and go break in, sort of thing, and then we have to basically down tools and fix them. But until then we can keep moving, right?" And then what happened was there would always be plausible deniability, right? "My test didn't make that go red. That change wasn't my change, that was Dave's change." 


Then we used to have somebody appointed on the team whose job was to essentially determine who had broken the tests and made them go red, and then that person would essentially apportion blame for the person that had to go and fix it, right? The fact that we actually had a step in our process which essentially was "somebody decides who's to blame for the tests being red today and then they have to go and fix them" implies that no one really cared about those tests, because they'd reached that point of being red the whole time. 


Really, the developers couldn't see any value. They constantly resisted writing these tests. And one thing I've learned over time is that, you know, if an idea actually really works, after a while your team will be quite happy behind it and you'll get some supporters of it. If your team is constantly pushing back saying, "We don't see what we're getting from this, I don't see the value in the tests that you're forcing me to write," then maybe you need to think again about what you're doing, okay?


And about a year after I gave this talk for the first time, DHH, the creator of Ruby on Rails, posted this blog post: *"TDD is dead. Long live testing."* And a lot of what his complaint was: "I am being forced by people to say that Rails is wrong, because Rails doesn't work the way they think test-friendly frameworks should work." So he said, "You know, I have the Active Record pattern in Rails. (For those who don't know Active Record, effectively the data persistence methods are actually on the object that you're persisting, so your user object has persistence methods on it.) I see people tell me this is no good because they can't cleanly separate their domain object models from basically their testing framework." And he said, "I just used to test... like write my unit tests and I talked to the database, and I'm now told that's wrong as well. And I said I used to like TDD when it first came out, but now I'm constantly being told I'm doing TDD wrong. Then I'm just giving up on TDD, you know? TDD as far as I'm concerned doesn't work. I'm still going to do some testing, but I'm quitting on TDD."


And of course, such a stir that Kent Beck kind of responded saying, you know, "RIP TDD." It was kind of a sarcastic post, which may not have been the best choice when Americans are amongst your audience, but he kind of said, you know, "Oh, well, here's all the things I'm not going to do now I don't have TDD anymore." So he's kind of defending TDD. And Martin Fowler actually produced this series of dialogues, mediated them between Kent and DHH, talking about TDD. DHH went and saw an original version of this talk I gave at NDC and was kind of saying, "Yeah, I think you and I are on the same lines." And they came to a pretty similar conclusion as the rest of this talk. So you might go away, if you like the ideas here, and go and see Kent and DHH having this dialogue mediated by Martin about what everyone thinks happened.


But the question is kind of, where did it all go wrong, right? What went wrong for TDD, and can we fix it? Can we change our ways? 


Anyone know... anyone ever hear of an English footballer called George Best? No, mostly... very, very famous in the 1970s. And the story goes that a bellboy at the hotel enters George's room in the hotel. George is lying on the bed naked with a couple of supermodels. The bellboy is wheeling in the caviar and champagne that he is going to give to George and these naked supermodels on the bed, and the bellboy looks down at George and says, "George, where did it all go wrong?" And so treat that sentence in that kind of light, right? I am a fan of TDD. What I want to talk about is, you know, where did it all go wrong?


Okay, so what I want to do really is go back to the beginning. Let's go back to the beginning of TDD, and I want to explain to you: the premise of my argument is TDD has answers to all the problems that we have been meeting with TDD. We have just ignored the fact that basically the original approach to TDD as described by Kent Beck works. Later patterns and practices we have loaded onto TDD, in many cases, were a mistake and have led us away from what was originally intended.


So I would recommend: who's read this book? (*Test-Driven Development: By Example*). Okay, many more of you practice TDD than have read this book. Who read some other book to learn about TDD? Who was told how to do TDD by somebody else that worked near them? Yeah, probably more of you, okay. 


I would recommend reading this book if you practice Test-Driven Development and you want to understand how it works. I recommend reading this book. When I began to struggle with TDD, I went back to read this book again. I'd read it originally, but I went back to read this book again. And one of the things I learned was that Kent had an enormous amount of wisdom about TDD, having practiced TDD for a large number of years before he wrote the book. And many of the problems I was encountering with TDD, Kent had already encountered, understood, and were discussed in this book. And this book, although it seems old, covers all sorts of things like mocking in TDD, which many people think are much later developments—they happened a lot later in the lifecycle; they're all in here, right? This book is probably the only wisdom you need, but I don't think I understood much of its wisdom until I'd been doing TDD for a certain amount of time, and I was then able to understand what Kent was trying to tell me, okay?


I think the most important thing to take away from an understanding of that book is this: **When you are practicing Test-Driven Development, do not test implementation details; test behaviors.** 

I'll explore that sentence more as we go through this talk. That is the key thing that people get wrong today.


So what happens is, in a classic modern TDD cycle, someone says, "I'm about to write a method on a class. I will write a test before I write that method, and that test will govern whether that method succeeds or fails." And so the trigger to writing a test in TDD practice is essentially adding a method to a class. 

That is the wrong thing! That is not the ethos of TDD. 


The trigger in TDD for creating a new test is that you have a requirement you want to implement. It is nothing more; it is nothing less, okay? There is something I want the software to do; I will write a test to do that, okay? I want basically the software to add an amount to this customer's bank account, essentially, right? I want the software to upload a new file to the user's collection of files, right? That's what drives a test in TDD. That's what you are writing tests for. You're not writing tests to say, "The `Foo` method should take the two numbers, multiply them, and return the result," okay? That is an implementation detail. There is no requirement you received—unless you're writing a calculator—no requirement you received that gave you that. The requirement you received was something else much higher up, right? And that is where you write the tests.


What you want to be doing is focusing on testing the API of your software, okay? When we talk about the API, we'll talk carefully about that in a second, but we mean: What is the contract your software has with the world? What does it offer to do for people who are basically consuming your software? That is the stable idea of your software. It will not change that rapidly. You may add new stuff to it, etc., but it's a stable requirement you have. How you implement that requirement is unstable: you may get better ideas over time, and you may wish to change it. But what your software offers to consumers is the stable contract, and that is what you should test.


We don't necessarily, by the way, mean by this an HTTP JSON API, right? So I'm conscious now that when I use the word "API," everyone imagines, you know, REST, HTTP, JSON, GraphQL, whatever. What we mean simply by an API really is just the publicly exposed surface area of your module generally that somebody else is consuming. So in, you know, the classical description of a module: a module has basically a set of exports—public-facing things that other consumers can call, that other software can call—and internal implementation details that you've hidden away. These are classic notions of information hiding, encapsulation, right? Your module has a facade, something on the outside that essentially represents and hides all those details away. And it is that external sort of exports that you are interested in testing, not the implementation details which tell you how you are implementing that, okay?


What happens is we should get some use cases or a story in an agile environment, and that says, "Here is the requirement I want you to implement. Here is the story, here is the use case, here is the thing that your software is supposed to be doing." And you write tests to say, "Right, I'm going to write tests to prove I can do that." And generally speaking, there are good models for this like Given-When-Then: Given that my account has £100 in it, when I add another £100 to it, then I have a total of £200, right? That is a kind of classical requirement you're trying to implement. Now I may end up using objects to represent accounts and that kind of thing, but what I'm doing is saying my requirement is the behavior of my account, perhaps when I add or remove things from it.


**The system under test is not a class.** Too much TDD practice focuses on essentially writing tests for classes, okay? The system under test is really the exports for a module, its facade. When people talk about this phrase "unit testing," at the time that people were basically coming up with this kind of idea, to them in classical testing paradigms, a unit test was a test of a module, right? It's kind of a black box where effectively you talk to basically what's exposed. 


Now, there's a certain line between classes and modules where effectively people can say, "Well, a class is a module—it has information hiding, encapsulation, etc." I appreciate that. But the definition of "unit" has become too much fixated on being a class. People talk about isolating that class and isolating the class's dependencies, right? Now, it leads into all sorts of problems, because then you start mocking heavily, which leads you into what I call over-specified tests where essentially you know all about the internal implementation details of that class through your mocks and how they're called, and essentially you've hidden no implementation details! 


So you want to focus at a much higher level. You want to say, "I have some public things that this module does; that is what I'm going to test. I have some implementation details, and I'm not going to test those, okay?" And we will explain how to get there. TDD explains exactly how to get there. This may sound complicated, but there is an easy route to get there, okay?


So obviously, you may actually be driving a class, because if you're writing say C# or Java, you can only drive classes. But that class may tend to be a facade or a command, etc., some kind of interface to the actual module underneath. You don't want to write tests to cover implementation details because if those change... you want to write tests for your stable contract, which is essentially your exports from your module, your API. 


Refactoring is the key step which enables you to achieve this goal of separating between things that are implementation details that you don't need to test and things that you do need to test. There is a get-out clause, and the get-out clause says sometimes it's helpful to probe implementation details that are difficult to understand—you can do that, but with test-afterwards; again, we'll come back to that. 


So a lot of these ideas... I want to give a bit of proxy to Dan North. Dan wrote a post about what he called Behavior-Driven Development. Now that's turned into something else really, though it has the same origin point, the same root. This is the root of that. When people talk about Behavior-Driven Development today, they mean something else entirely; we'll talk about that later. But in his original post, Dan said, "I've often found when I go to basically implement TDD in places, people are very confused about how to do TDD because of the word 'test'. The word 'test' throws them. And if I change that word to 'behavior-driven'—developing your system using behaviors—people do it correctly." This insight was very old, you know, around 2005, etc. And he's saying exactly what really I'm going to tell you in this talk: people misunderstood what Kent was talking about, probably because Kent used classical testing terminology like "test", "unit test". He mentions it just a couple of times in the book; generally, the phrase Kent Beck uses is "developer test", right? Not unit testing—developer tests. 


But Dan basically effectively recognized the same problem I'm recognizing here, which is people were driving tests by saying, "I need to test this method." They should have been saying, "I have a behavior in the system, and I want to use code to basically prove that that behavior works in the system, okay?"


So in case you think that's just a Dan thing, Kent wasn't talking about this... Kent Beck originally, Kent talked about behaviors: "Behaviors would demonstrate that code would compute a report correctly." So he's saying the object of TDD is to test behaviors in the system, okay? 


Here is a classic thing that you are testing: "Add amounts in two different currencies, convert the results with a given set of exchange rates." We are not testing a method on a class; we are testing a behavior of the system that we can add two different amounts in two different currencies and then see basically the result in one currency. "We are multiplying a price per share by number of shares to receive an amount." Okay? The behavior is the requirement. These are TDD requirements as described by Kent Beck originally. That is what your test should be looking at, not "I need a method on my class for exchange rates." That's nothing... that's not... nothing to do with TDD.


So the idea behind Test-Driven Development has always been that we are essentially saying: What should the API of our code look like? If I am an external consumer of this code and I come to call it, what kind of API would I find useful? You're testing as much the usability of the API—how you interact with it—as you are essentially designing that public-facing API as anything else, okay? So you tell yourself a story about how this API should look if you were a consumer. The test becomes the first consumer of your code that essentially says what should this look like and how should it work, right? That is the objective of Test-Driven Development.


So Kent uses "unit test" in a very specific way, right? And simply for him, when he uses the phrase, all he means is that you should be able to run all your tests together as one suite without one test being run impacting the other test. 

**The unit of isolation is the test.** 

And a lot of mistakes have been made because people believe the unit of isolation is the class under test. It is not, okay? 


So all these people mocking because they say, "Well, you've got to have the class isolated when you test it. These all have to be mocks. You can't interact with anything else. You've got to mock all the class's dependencies, right? Because otherwise it's not a proper unit test." You are wrong! The unit of isolation is the test, not the thing under test.


One of the problems is, you know, people were using definitions from classical testing. And the thing is essentially unit tests referred to this idea of a module, right? This is why it's a confusing phrase. Because when people talked about unit test, they were saying: take a module—so say, in .NET terms, an assembly or something, right?—take that and then essentially test it on its own. Then an integration test is I've got a couple of modules talking to each other, and then essentially a system test is talking... all modules talking to each other. So when Kent Beck used this phrase "unit test", what he was really talking about was a module being tested, okay? It's an unfortunate phrase, because people then said, "Well, the unit is a class," and came up with a whole set of practices which basically have been negative towards our TDD practice.


You may choose to avoid using certain interactions in testing, right? Databases, file systems, those kinds of things, right? But the reasons Kent gives for that are slightly different from the ones you think. It's not about isolation—it's about isolation only in the sense that if I touch a database, one test may impact another test because if I add some rows in the first test, when I read them in the second test, that may be impacted basically by the order they run in. So I have to make sure that between these two tests, I'm isolated from that kind of problem. Which may mean I choose to use something like an in-memory database and we tear that down rapidly and provide a fresh one for each test, right? But it's the isolation of the test that is driving that, not the isolation of the code under test.


Same with file systems, everything else. And the other issue is speed, right? Tests have to run fast; otherwise developers don't get fast feedback. So really, I want my test suite to run in a couple of minutes, and it's quite possible if I'm talking to real hard dependencies that becomes hard. So I may wish to essentially swap those out for the purpose of speed. But those are the reasons for it; there are no others. If there is no shared fixture problem, it's perfectly fine in a unit test to talk to a database or a file system. And that happens all the time in Kent Beck's testing. So people telling you, "You can't talk to a database in a unit test," they are wrong, okay?


Yeah, okay. Focusing on the methods creates tests that are hard to maintain, okay? Because essentially, when I look at that test, what I see is somebody describing how you interact with a method. That's why we get this problem when we come back to our tests: they're very hard for us to understand because we're kind of going, "Well, what does this test doing? Well, we've got a method in isolation, so it's very hard to understand what it's doing." If the test tests a behavior, it's very easy to understand what that test is doing, because it says, "This behavior is asserted to be true in the system." Now that test is red because I've changed another part of my code, I can either say, "Well, that should still be true," or perhaps, "No, that isn't true anymore. That behavior's been modified now by the new requirement, which means this behavior has also changed and I've just seen an impact of that behavioral change."


And when we have implementation details exposed—in other words, essentially all those things that make up how our API actually works—that's when you get this problem effectively of being hard to change, particularly if you use mocks. So when I use mocks effectively, when I say that's just a class under test and I know about all the calls out to mocks, the problem is when I change something that changes those, all the basically assertions break. And if I essentially therefore know too much about the implementation details, then I've coupled my tests to my implementation, and I can't change one without changing the other.


All right, so let's talk about how we get there, because the way we do this is really simple, and it's Red-Green-Refactor. 

Who thinks they understand Red-Green-Refactor? Okay, great. 

Who does Red? How many of you make sure that essentially... okay, very few of you. 

Who does Green? Probably most of you. 

Who then always does the refactoring step? Some, okay. 


Refactoring is the key thing, right? 

The test that doesn't work: okay, this is important. There are no tests for tests. I have to demonstrate that my test will fail in the absence of correct implementation, right? I've stumbled across this a few times, had the problem where essentially I had a test that would always pass. And so, you know, essentially until I got that red phase, I wouldn't prove anything by writing my code and having it pass again, because it was always going to pass anyway.


Green: Make the test work quickly, committing any sins basically necessary in the process. This is really important. What I'm doing is making the test pass by the quickest process possible, okay? A good solution here: cut and paste code from Stack Overflow, right? This is where the duct-tape programmer wins, because you're afraid to do that and he's not afraid to do that, okay? Now, make good code... but it's three steps: write a test, make it compile and see that it fails, make it run. 


And then the third step is refactoring, okay? 

When you write your implementation of your test and you go to your green phase, what you are seeking to do is understand how to solve the problem. You want to basically, as fast as possible, get something that satisfies the behavior: "I want to add two currencies in different amounts, and I want to basically look at the end result, right, in one currency." What I want to do as quickly as possible is figure out how to do that in my code. Just write a good transaction script: one line, with one line after another. Don't try and basically put it into lovely classes; don't try and essentially put patterns in there. Just write line after line of code until it works, okay? And that code can be dodgy if you want it to be—probably should be! 


Your goal is to, as fast as possible, get that test to go green because you have managed to understand how you will implement the requirements. At this point, you are trying to determine how you will implement the requirements, and you can figure out the answer to that by going green. Go to Stack Overflow, copy some code, right? You are explicitly enjoined in this step to write non-well-engineered code. You want to be the duct-tape programmer! You want to be that really annoying guy on your team who writes bad code. Emulate him! Do what he is doing to get it out the door fast. You will now be keeping up with him. You've added the test, but he has probably done it manually a few times; you are moving at the same rate, okay?


As Kent Beck says: **For this brief moment, speed trumps design.** 

I want to get there as fast as I can, okay? Kent's point is essentially you can't do two things at once easily: you can't both understand the solution to the problem and engineer the code, right? One of two things will go wrong: you will either over-engineer your solution beyond what the test actually requires, or you'll get some kind of analysis paralysis, right? But the first thing you do is solve the problem with the code.


Okay, so when we're getting our clean code... okay, so we believe in clean code; we don't want to have duct-tape programmer code all day long. And we're going to do that in the next phase: the refactoring step. And to produce clean code... so for Kent, basically the key driver is duplication, right? You see duplication, you remove that; that starts to re-engineer your code well. And that works there. 


You can add into that, I think, Fowler—Fowler's book *Refactoring*, I'll talk about that a bit later: look for code smells, okay? Look for smells in your code that indicate there are better ways of doing this. 

And Joshua Kerievsky's point is: if you want to think about using design patterns, this refactoring step is when you identify them. Because now you have a clear idea of what your solution is, you can tell whether a pattern is appropriate or essentially whether you're using patterns inappropriately as a solution, okay?


When you refactor—and this is the really important part—**you do not write any new tests.** Refactoring is a process of safe moves that let you essentially change the design of the code; they do not change its behavior. Your behavior is covered by the original test. The original test tells me I'm adding two amounts in two different currencies and getting the final result. The code I'm implementing produces that result. If I use safe refactoring steps to get there, it doesn't matter how much I change implementation details; I am still covered by my original test. 


Code coverage tools help you here, because you don't want to introduce, for example, new conditionals, speculative code at this point. When you do that, you need to have new tests to cover those new behaviors, because generally you've got a conditional because you have a new behavior that you're trying to produce. But do not write new tests at this point!


This results in two things: 

One, you will write less tests, so you will move faster. The refactoring step shouldn't take you so long that you fall miles behind this duct-tape programmer, right? You will keep up in the way that the lean software guys and your managers want you to, because you are not saying, "Oh, I'm creating a new class here, I'm doing extract class because I believe that essentially these things should be a class in my implementation." Don't write a test for that class, right? 


If you have a language that supports visibility—for example, Java or C#, and you've got things like protected, internal, private, public, right?—in C#, for example, most need to be internal and never exposed. And as such, the test should not be coupled to them. You should be at liberty to change their implementation whenever you like, and no test will fail because no test is coupled to their implementation. 


And this is key to understanding this, right? In the refactoring step, do not write new tests! You're now introducing a couple of classes. If you basically keep thinking, "Mmm, I need to introduce a public class," the only smell you may be looking at is that essentially in your solution there is some new thing that is also part of your API, okay? That will itself be exposed as part of the API of your module. In that case, that's probably an appropriate time to think about a mock for use in your test, and later go and implement that other part of your API that you're currently missing, right? But focus initially on trying to avoid writing these extra tests.


It says, right: dependencies are the key problem in software, right? Coupling is a worse enemy for you than DRY. Coupling is what kills all software. So struggle to decouple your tests from your implementation details. Focus on your public contract, your stable API, and leave your implementation details free of tests, and avoid heavy mocking. 


This allows you to meet the promise of refactoring: you will refactor your code and your tests won't break, right? I've seen this work. Since I changed this practice, it's been very easy for us to refactor implementation details, to figure out new ways of doing things. We changed the whole way we did, for example, modeling an aggregate without any tests breaking, because the behavior hadn't changed; our choice about how to implement it had changed. So test behaviors, don't test implementations. (These mics are annoying if you wear glasses, right? Okay.)


Keep internals internal. Don't make these public to test. If you find yourself making things public in order to get them under test by implementation details, ask yourself if that genuinely is a part of the contract of your module. If it's not, don't! It's just implementation detail; why did you surface it for testing? If you are in .NET land and you use the assembly attribute `[InternalsVisibleTo]` and list your test assembly so that you can get your internals tested, do not do that, right? Your internals should not be visible to the test because they're not under test, okay? So preserve your information hiding; have a public API and internals.


Right, this is Fowler's book *Refactoring*. Actually came out before Kent's book on TDD, and it talks really about code smells and says here are some safe steps to be able to effectively refactor basically your code when you see these smells. And one of the key things to understand is, you know, very explicit, okay: we are changing implementation details, not public interface. 


Anyone ever use the DevExpress tools as opposed to ReSharper's tools for refactoring? So DevExpress had a set of tools which basically had a code plug-in for Visual Studio, and they used to offer refactoring. And DevExpress's tools... some of the guys that worked on them were very clear, right: you cannot refactor a public method on a public class. They wouldn't let you do it! Because that's a change—that's not refactoring, because you don't know the impact of that. Now later they said, "Okay, you can press the button marked 'Unsafe Refactoring'," because everybody else lets you do it and we're being killed in the marketplace. But also because they said, "Well, you may say to me the only consumers of this public API are within the span of the solution, and therefore essentially I understand who all the callers are and I can then refactor them." So if I'm doing a rename, I know who I'm going to break. But the goal was originally: don't do this, right? Don't change... refactoring is not changing your public interface; it's changing how you chose to implement that. It is a safe, small change. There are recipes to do them, that's why tools automate them, and they're safe in the sense that your code is not basically going to... it's not going to negatively be impacted if you extract a method; it's a safe change, okay?


Here's the list of code smells. It's not really what this talk is about, but those are the kind of things that you're looking for, and those are the things that drive you in your refactoring step to do things like extract a class out of this transaction script body, extract a method out of this body, right? Look at basically where you've got Feature Envy, whether you're doing some kind of dot syntax, etc., right?


And Joshua Kerievsky in his book (*Refactoring to Patterns*) said there's a problem with patterns, and the problem with patterns is people are getting pattern-happy. They see patterns everywhere and start applying them, and in many cases what they've done is applied them in the wrong context or made their software over-complicated when it didn't need them. What Joshua Kerievsky said is: the ideal step for applying patterns is the refactoring step. Because at this point, I know what the solution is—I've written one! It's an awful solution, because I was told to make an awful solution! And at this point now, I can see I can clean it up by making a pattern. I can see that actually this is... which is the Command pattern, I need that here; this is the Template Method pattern, I need that here. But I know that now because I've already implemented it once badly. And even in Kent Beck's original book, he identifies patterns you may want to refactor to, so this is part of original TDD, it's not new stuff, okay?


So what about Ports and Adapters? 

This is unfortunately the common way that testing occurs in a lot of organizations:

Before we release the software, there's a big, big wave of manual testing going on. Everyone's getting into the system, running manual scripts, etc., doing some kind of conference releasing, right? 

Then below that, there's some sort of scientist who has big suites of Selenium tests that they run against their software. Not so bad? Funny of you, okay. 

And then we have integration tests, and at the bottom we have a few developer tests, right? 

This is a bit of a disaster! Why? 


Well, first of all, this manual testing is really expensive, unrepeatable, so you really want to automate your way out of that situation. 

This UI testing one is very problematic, and I'll tell you why: if I change my UI to give it a nice classy new look and feel that's much more, you know, 2017 rather than 2008, right, I may not be changing any other behavior of my software whatsoever. They do exactly the same thing that they did yesterday; I've just cleaned up the UI. But these tools all break, okay? But I'm driving them to test the behavior of my system, but the tools all break because I changed the way that my widgets work, okay? So they're very fragile, brittle, and slow. They also run really slow—quite often you run them overnight, and then you get the blame game in the morning: someone's job in the organization is to determine who is to blame for last night's tests failing, right? If you find yourself in the situation where you have somebody whose job every morning is to determine who's to blame for tests failing, you are doing it wrong! I know that because I've done that, and that's wrong, okay? All the mistakes I'm telling you about, right, I've made those! Don't worry, I'm not going at you because I think I'm smarter than you; I only know this because I've made every single one of these mistakes, right?


So what you want is what's called the Testing Pyramid: the majority of your code should be tested by your developer tests; a small amount of code up here should test your widgets work correctly; and in between, there's a certain amount of integration testing where the two get hooked together. There's a bit of a scam in the middle; I'll come back to that later, okay.


So this is a Hexagonal Architecture. It's gained popularity recently. People like Bob Martin called this the Clean Architecture; some people call it an Onion Architecture. There are various names for this architecture, various attempts at co-discovery of the same thing, all right? 


What happens? 

In the middle, we essentially have Plain Old C# Objects, Plain Old Java Objects, Plain Old Python Objects—anyone here ever heard that phrase "POCO"? That's also slang for the police in the UK, so be careful about using that word, but yeah! But these things are devoid of any technology concerns, so they don't know how to talk to the database, they don't know how to talk to the web; they're just basically completely plain old objects without any technology concerns.


On the outside, we've what's called the adapter layer. The adapter layer knows how to talk to the outside world; it has all the technology concerns. So our primary ones over here: things like basically an HTTP adapter, so something like basically your web framework—Web API, Django, Sinatra, effectively, or Flask, MVC, etc.—or a GUI adapter, so like Web Forms, or, you know, JavaScript, like that, right? And they talk basically to your domain. Over here, we've basically got a database that your domain uses to persist itself; over here, we've got essentially the idea of things like sending messages, things like RabbitMQ, or essentially that kind of thing, right? Or emails.


In between is what's called the ports layer. So the ports layer literally is like a port in an operating system: it is a door. It's the door which essentially allows one boundary to be crossed. Here, it's the door that essentially says, "Well, my HTTP code needs to talk to my domain; it would go through this door," some kind of facade layer which says we talk to it. Over here, it's saying, "My domain code needs to basically persist to a database, so I have a door that lets me talk to databases," some kind of repository model perhaps.


Now, okay, I've got a few five minutes more, and I'll just run through these quickly, right? Apologies for running over at the start, okay.


So the port here is the ideal point to do your... for your tests! The tests talk to that port layer, not to the internals here. Because the port is how everything communicates with your code, and also how everything your code raises—like events, or calls the database, okay? So the port is the point for you to make your calls. 

Integration tests are really just checking that essentially you can talk to the database correctly, you're configured correctly on this adapter. Don't test adapters themselves; they're usually third-party code. And a system test just checks that you can call from the outside in, right? And that is essentially how you do it, okay?


Make two quick points:

I'll talk very briefly about gears. So you may say that's all great, and testing the API is fine, but I sometimes like TDD to help me understand how something's working, and I want to basically drill down using my TDD into the details, okay? 


So Kent Beck has this notion of gears. How it works: it's notionally five gears. 

In fourth gear, it's the practice you just described: you're testing the public API, and you're getting basically the full script out, and you're refactoring that, right? 

In fifth gear, it's kind of like the implementation is very obvious, so actually that refactoring step—when I write my code, rather than write the sinful code, I just write the correct answer, so it's a regular refactoring step. Be careful here: if basically you find that your coverage starts to blow out when you put your tests, shift down to fourth, because you're probably writing too much speculative code you haven't got tests for. 

If what I find though is that I'm saying, "I don't know how to go green, I'm not really sure how to implement this, and I want to feel my way with the aid of tests to probe me a bit more," then shift gears down and write some tests to help you do that. You could also use things like a REPL—REPLs in languages like Python, etc.—for this technique of just feeling your way a bit. But feel your way, but at the end of that process, **delete the tests!** The tests helped guide you, but now they're a burden to somebody else coming along. The person coming along who wants to basically change the implementation details probably means that your way of doing it, they're changing it altogether. Your tests are something they don't understand; we're just protecting your implementation details. Throw them away, and let the next person coming along write tests if they need to check the implementation, okay?


ATDD, very quickly:

Kent Beck originally said don't do it, because it has two problems: customers are not interested, and it's very expensive to do, right? He later, in 2011 or so, had a book about ATDD in his series: "Maybe I was wrong, maybe actually you need this because programmer tests are a problem, they don't focus on the gist of the system." 

James Shore, who worked with Fit, who was one of the guys that wrote one of these tools, says, "I don't do it anymore." They said the reason, in fact, over the years, that customers essentially don't care, and they're a significant maintenance burden, right? So all these tools like Fit, Gherkin, that kind of thing, right? He just says actually, if I have customers again that help developers drive the developer tests, so the unit tests are informed by the developers talking to the customer and write the appropriate unit tests... 


Over time, I used to believe in these tests; I used to think that was worth writing, you know, Fit and Gherkin tests, Cucumber, SpecFlow tests, because I thought it helped keep developers honest, we'd know what "done done" was. But over time, I've come around to realize that probably what we were solving was the problem of how we were doing our unit testing. Because our unit testing was focusing on methods on classes, we had nothing that told us about behaviors! But what we actually needed to do is shift our TDD practice to be about behaviors, and then we can drop ATDD altogether. Just write developer tests in the same language you're using to implement, and drop this horrible translation layer between natural language and your codebase.


Behavior-Driven Development:

I mentioned briefly Dan North, basically the guy who originated it, had this original post which basically describes testing to be about behavior. It is now a full-scale agile methodology; you may be interested in that. But this talk is not about Behavior-Driven Development as a methodology, right? It's taking that original kernel idea that Dan had of saying: test behaviors, don't test methods on classes.


Mocks, briefly:

Mocks are useful when a resource is expensive to create and represents a shared fixture problem. But the problem with mock objects essentially is people have used them to isolate classes. Don't do that! And especially don't do this thing where effectively you say, "I'm going to make sure junior developers have done the right thing by mocking. I force them to use mocks so I can see how they implemented the method." It just couples your test to your software, means you can't change it; it's never worth it. They over-specify software.


Generally speaking, IoC containers are overused. Try and use them as little as possible. Only use them because your framework demands that you actually do it because they don't understand how to create objects that you have. If your framework was good, it would actually ask you to give it factories, not an IoC container. And you might implement your framework factory in terms of an IoC container. If you find yourself implementing an IoC registration in your test, something has gone horribly wrong, okay? Bit like ORMs, we overused IoC containers; please avoid that, okay.


Right, quick summary:

- Reason to test is behavior, and basically because we have a new behavior.

- Write dirty code to get green, and refactor.

- No need to test refactored internals and privates; they're open season against tests.

- Write on a port in Hexagonal Architecture.

- Add integration tests only really to the extent you're covering basically how ports are implemented.

- Don't mock internals; ports are adapters, not one little bit missing out of defended code.


Okay, I ran out of time for questions, I'm sorry. I'm going to let you all go. Please come up to the front if you want to ask a question at the end, or come and find me during the conference. I'd love to talk about this stuff. You can see this talk online; versions have been given in places like NDC, so if you want to share with your colleagues you can go and dig out the video, and I think these guys are videoing it too, so you can spread the word with your colleagues. It's worth going and looking at that conversation between Kent Beck and DHH, mediated by Martin Fowler as well, for a similar conversation about how it's all gone wrong. 


But thanks very much! Thank you. [Applause]


n5321 | 2026年10月7日 19:47

Full Walkthrough: Workflow for AI Coding — Matt Pocock

>> Yeah, we're good.

Okay, folks. We're at capacity. Let's kick off. I don't want you waiting here for 25 more minutes before we some arbitrary deadline. So, welcome. My name's Matt, I'm a teacher, and I suppose now I teach AI.

Um. We have a link up here, if you've not already been to this, which is has the exercises for the um stuff we're going to do today. This is going to be around 2 hours, so we might just sort of kick off 2 hours from now. Is that all right, Mike?

>> Yeah, perfect.

Um and the theory behind this talk, or at least the thesis under which I've been operating for the last kind of 6 months or so, is that we all think that AI is a new paradigm, right? AI is obviously changing a lot of things. You guys are obviously interested in this, and that's why you've come to this talk. And I feel that when we talk about AI being a new paradigm, we forget that actually software engineering fundamentals, the stuff that's really crucial to working with humans, also works super well with AI. And this is what my keynote is on tomorrow, really. I'm going to sort of be fleshing that out a lot more. And in this workshop, I'm hopefully going to be able to direct your attention to those things, and uh hopefully show you that I'm right. But we'll see.

Um can I get a quick heads-up first? How many of you guys um are coding have ever coded with AI? Raise your hand if you've ever coded with AI. Perfect. Okay. Uh keep your hand raised. Uh let's all uh share those armpits with the world. Um how many of you code every day with AI? Cool. Okay. Uh right, keep your hand raised if you've ever been frustrated with AI.

Okay, very good. You can put your hands down. Thank you for that show of obedience. I really appreciate that. And we are also being live-streamed to the Gilgood room as well. I've not uh Did we send someone up to the Gilgood room to just check they're okay? Don't know. But I see you, and there is a way that you can participate, which is we have the um a Q&A. We're going to be doing kind of have a sort of hatred of Q&As cuz they're not very democratic. They're mostly the sort of um most talkative people get to um get to participate and share. And so, we're going to be going through this um Q&A here. So, why do we have to wait till 3:45? The room is packed, the doors are closed. 100% agree. And so, if you want to uh ask a question, we're going to be I would like you to pile into this async, and then we can vote on each other's questions, and hopefully get the best questions surfaced so the for the entire room to enjoy.

So, I want to talk about first the kind of weird constraints that LLMs have. And those weird constraints are sort of what we have to base a lot of our work around. Now, there's a guy called Dex Hardy who runs a company called Human Layer, and he came up with this idea, which is that when you're working with LLMs, they have a smart zone and a dumb zone.

When you're first kind of like working with an LLM, and it's like you've just started a new conversation, you start from nothing, that's when the LLM is going to do its best work. Because in that situation, the attention relationships are the least strained. Every time you add a token to an LLM, it's kind of like you're adding a team to a football league. You think of the number of matches that get added every time you add a team to a football league, it just goes it scales quadratically. And that's because you have attention relationships going from essentially each token to the other that are positional and the sort of meaning of the individual token.

And so, this means that by around sort of 40% or around I would say around 100K is kind of my new marker for this. Cuz it doesn't matter whether you're using 1 million uh context window or 200K, it's always going to be about this. It starts to just get dumber. So, as you continually keep adding stuff to the same context window, it just gets dumber and dumber until it's making kind of stupid decisions. Raise your hand if that feels familiar to you.

Yeah, cool. So, this means that we kind of want to size our tasks in a way that sticks within the smart zone. Right? We don't want the AI to bite off more than it can chew. This goes back to old advice like Martin Fowler in refactoring. Uh like uh the pragmatic programmer talks about this. Don't bite off more than you can chew. Keep your tasks small so that you as a developer, a human developer, don't freak out and don't start acting and going into the dumb zone.

But how do you tackle big tasks? How do you take a large task like I don't know, cloning a company or something, or just doing something crazy, and how do you break it into small tasks so they all fit into the dumb zone? One way, of course, you could do is I mean, kind of what the AI companies maybe want you to do, or the natural way of doing it is just keep going and going and going, you end up in the dumb zone, charging you tons of tokens per request. You then compact back down. We'll talk about compacting properly in a minute. And you keep going, keep going, keep going, compact back down, keep going, keep going, keep going. And I think that's doesn't really work very well because the more sediment I we'll talk about that in a minute.

So, the theory here is then, and this is what I was doing for a while, is I would use these kind of um multi-phase plans. Where I would say, "Okay, we have this sort of number four thing here, this large large task. Let's break it down into small sections so that we can then kind of chunk it up and do each little bit of work in the smart zone." Raise your hand if you've ever used a multi-phase plan before.

Yeah, really common practice, right? This is kind of how we've been doing it. Certainly, this is how I was doing it up until December last year, really. And any developer worth their salt will look at this and go, "This is a loop." Right? This is a loop. We've just got phase one, phase two, phase three, phase four. Why don't we just have phase N? Right? Phase N. Where we essentially just say, "Okay, we have, let's say, a plan operating in the background, and then we just loop over the top of it, and we go through until it's complete."

And this is where um Raise your hand if you've heard of Ralph Wiggum as a software practice.

Okay, cool. Raise your hand if you've not heard of Ralph Wiggum as a software practice, actually. That's more like it. Okay. So, there's this idea called Ralph Wiggum, uh which is kind of um sort of based on this, which is essentially all you need to do is sort of specify the end of the journey, where you just say, "Okay, we create a PRD, a product requirements document, to say, 'Whoa, okay, let's describe where we're going.'" And then we just say to the AI, "Just make a small change. Make a small change that gets us closer and closer to that." And Ralph works okay, but I prefer a little bit more structure. So, that's kind of where we got to in terms of thinking about the smart zone, and that's kind of where I want you to first start thinking about here.

Another weird constraint of LLMs is LLMs are kind of like the guy from Memento, right? They just continually forget. They could just keep resetting back to the base state. Let me pull up this diagram. I sort of I I I really should use slides, but I just prefer just like randomly scrolling around a uh infinite uh TL draw canvas. Thank you, Steve.

Um. So, let's say another concept I want you to have is that every session with an LLM kind of goes through the same stages. You have, first of all, the system prompt here. This gray box here is essentially the stuff that's always in your context. You want this to be as small as possible. Cuz if you have a ton of stuff in here, if you have 250K tokens, like I have seen people put in there, then that you're just going to go straight into the dumb zone without even being able to do anything. So, you want this to be tiny.

>>[snorts]

You then go into a kind of exploratory phase. This blue sort of where the coding agent is going out and exploring the code base. Then you go into implementation. And then you go into testing. And sort of making sure that it works, running your feedback loops and things like this. Raise your hand if that feels familiar based on what you've done. Yeah. Sort of the like the the main cornerstones of any session. And when you clear the context, you go right back to the system prompt. Oof, you go right back there. So, you delete everything that's come before.

And raise your hand if you've heard of compacting, as well. Yeah, okay. There are some people who've not heard of compacting. So, let's just quickly show what that means. For instance, I've just been having a little chat with my LLM. Uh I want to make sure we sort of, you know, just cover the basics so we're all sort of on the same wavelength here. I've just been having a chat with my LLM. I've been talking about a thing that I want to build. How's the font size? Should I bump it up? Folks in the back? Bump. Bump. Bump. Bump. Bump. Oh. I'm using Claude Code for this session, but you don't need to use Claude Code. Um so, I've been having a chat with the LLM, just sort of planning out what I'm going to do next. It's asking me a bunch of questions, and I can I highly recommend you do this. There's this tiny little status line here that tells me how many tokens I'm using, the exact number of tokens I'm using. Um I have a article on my website AI Hero if you want to copy this. This is Oh, wow, that is that shakes, doesn't it? Um this is essential information on every coding session cuz you need to know exactly how many tokens you're using so that you know how close you are to the dumb zone. Absolutely essential. And so let's watch it. So I've got two options. I can either clear wrong and go back to nothing or I can compact. And when I compact then it's going to squeeze all of that conversation, which admittedly isn't very much, into a much smaller space. And this in diagram terms kind of looks like this. Where you take all of the information from the session and you essentially create a history out of it, a written record of what happened. And devs love compacting for some reason, but I hate it. I much prefer my AI to behave like uh the guy from Memento because this state is always the same. Always the same every time you do it. You clear and you go back to the beginning. And so if you're able to do that and you're able to optimize for that then you're in a great spot.

So that's kind of the two things I want you to think about with LLMs, the two constraints that we're working with. They have a smart zone and a dumb zone and they're like the guy from Memento. So let's take a look at the first exercise. And I'm while I'm doing this, the way I want this to work is I'm going to sort of show you how um I'm going to be sort of walking through it up here and I want you folks to be kind of like tapping away and doing things as well. So that was just a little lecture bit. Let's now actually get and do some coding. For anyone who arrived late or anyone in the Gilgud room uh go to this link this link up here to see the exercises and clone the repo. You absolutely do not have to, you can just watch me do it if you fancy it. But let's go there myself and let's see what exercises await us.

So essentially I've built a um this is from my course. This is a uh a course management platform essentially, a kind of CMS for instructors, for students, and this is what we're going to be building a feature in. So I'm going to take you from essentially the idea for the feature all the way up to building a PRD for the feature, all the way up to implementing the feature. And hopefully you can take inspiration from this process and use it in your own work.

So uh let's kick off. So we're going to start by using a a skill which is very close to my heart. It's the grill me skill. And this grill me skill is wonderfully small wonderfully tiny and it helps prevent one of I think the main issues when you're working with an AI, which is misalignments. The uh the sort of silent idea that I'm talking against here, that I'm arguing against, is the specs to code movement. Has anyone heard of the specs to code movement? Raise your hand. It's not really a movement I suppose, it's just sort of people saying specs to code. Um what it is is people say, "Okay, you can write a program or you want to build an app the best way to build that app is to take some specifications so to write some sort of like document and then turn that document into code." So they just turn it into code. How do you do that? You pass it to AI. If there's something wrong with the resulting code, you don't look at the code, you look back at the specs. You change the specs and you sort of just keep going like this. This is kind of like vibe coding by another name where you're essentially ignoring the code. You don't need to worry about the code. You just sort of keep editing the specs and eventually you just keep going. And I tried this. I really tried it. And it sucks. It doesn't work. Because you need to keep a handle on the code. You need to understand what's in it. You need to shape it because the code is your battleground. And so this is again is where we're going. Let's let's get some exercises.

So what I'd like you to do is go to this page, the the grill me skill. And inside the repo here we have a slack message from our pal. Uh where is it? It's in the root of the repo and it's under bur bur bur bur Oh, where is it? Mhm mhm client brief.md. It's a slack message from Sarah Chen. For some reason the Claude always chooses Sarah Chen as the name. I don't know why. Um it's saying that in cadence, our um course platform, our retention numbers are not great. Students sign up to a few lessons then they drop off. I'd love to add some gamification to the platform. And so when you're presented with an idea like this, you need to find some way of turning it into reality. Let's say Sarah Chen is your client, you're on a tight budget, you need to get this done fast. How do you go and do it?

Um raise your hand if you would um enter plan mode when you're doing this. Anyone a big user of plan mode? Yep. Um let's actually shout out quickly any other ideas about what you would do with this or any Raise your hand if you what what would be your first port of call?

>> Yep. Ask for more info.

Sorry? Ask for more info to verify what is the purpose and where our current standing is. Yes, exactly. Let's imagine that Sarah Chen's gone on holiday, you have no idea, right? Uh she's just posted this thing, you need to action it before you go. Well, my first port of call is I go for this particular skill. I'm going to clear my context. I'm going to uh get rid of you, you don't need to be there. And I'm going to say um I'm going to invoke a skill which is the grill me skill. Let's quickly check. Raise your hands if you don't know what this is.

Cool. Oh, sorry sorry. Let me be more specific. Raise your hands if you don't know what I'm doing here when I uh do a forward slash and then type something. Anyone Everyone kind of understand what that is? I'm invoking a skill. I'm invoking the grill me skill. And what I'm going to do is I'm going to say grill me and I'm going to pass in the client brief. So now the LLM really has only a couple of things here. It just has the skill and it has the description of what I want to do. And this is virtually how I start every piece of work with AI. And while it's exploring the code base I'm just going to show you what the grill me skill does.

So this is inside the repo so you can check it out. It's extremely short. "Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the decision tree resolving dependencies one by one. For each question provide your recommended answer. Ask the questions one at a time uh blah blah blah." What this does and what I noticed when I was working with AI, especially in plan mode actually is it would really eagerly try to produce a plan for me. It would say, "Okay, I think I've got enough. I'm just going to poof plan plan." And what I found was that I was really trying to find the words for this, for for what I wanted instead of that. And Frederick P. Brooks in The Design of Design, he has a great quote uh talking about the design concept. When you're working on something new with someone when you're uh all trying to build something together then there's this shared idea that's shared between all participants and that is the design concept. And that's what I realized I needed with Claude. I needed I needed to reach a shared understanding. need an asset, I didn't need a plan, I needed to be on the same wavelength as the AI, as my agent. And this is an extremely effective way of doing it. So hopefully Here we go. Nice. It has done its exploration first of all. It's invoked a sub agent which spent 97 93.7k tokens on Opus. Um and it's asked me the first question. Cool. We can see that even though the sub agent burned a a ton of tokens I haven't actually um uh increased my token usage that much. Raise your hand if you don't know what sub agents are. It's important question. Everyone kind of clear what sub agents are? Okay, I'll give a brief definition. Which is that this this sub agents thing here, this explore sub agent it has essentially gone and called another LLM which has an isolated context window. And then that LLM has reported a summary back. So a sub agent is kind of like a delegation. You're delegating a task to a sub agent. It goes eagerly does all the thing, explores a ton of stuff and then just drip feeds the important stuff back up to the orchestrator agent. To the parent agent. So okay. So hopefully you guys have seen the same thing. It's done an explore. And we now have our first question. Points economy. What actions earn points and how much? Ooh, okay. At this point you can ask it by the way questions to um deepen your understanding of the repo. I obviously know this repo really well cuz I wrote it, but you might not um know what's going on. So let's say my recommendation, keep it simple, two point sources to start. What's so nice about this is that not only does it give us a question that kind of aligns us here, we get a recommendation too. And often what I'll find is the AI's recommendations are really good. And so I'll just say skip video watch events, they're noisy and gameable. I agree. Sarah's asked we'll keep the lessons in the bread and butter. Yeah. Looks good, pal.

>> [snorts]

Now what I usually do is I usually dictate to the AI. I'm usually actually chatting to the AI instead of uh typing here, but uh this is a relatively new laptop and I couldn't get my dictation software working on it um because Windows is crap. Um So, should points be retroactive? There are existing lesson progress records with completion at timestamps. This is a really nasty question, right? Should we actually go back and backfill all of the lesson progress events? This is a kind of question that you need to be aligned on if you're going to fulfill the feature properly. This is not something I considered and Sarah Chen certainly didn't consider. Do I want it to be retroactive? Hmm. Let's actually do a vote inside here. Should we go back and backfill all the records? Raise your hand if you think we should backfill all the records. Raise your hand if you think we shouldn't backfill all the records. There are a lot of fence-sitters in the room. I'm going to say you know, this is the kind of discussion you're sort of having with the AI. You're getting further aligned. Yes, I'm just going to go with his recommendation cuz I'm lazy. Notice too how I'm able to keep in the loop here with AI. I'm not you know, it's it's pinging me these questions pretty quickly. I'm not having to go off and check Twitter or something. Levels. What's the progression curve? Yeah, that looks about right. For instance, yes, okay. So hopefully you should be able to go and um kind of work through this with the AI.

>> [clears throat]

And essentially try to reach an alignment. And this grill me skill, this can last a long time. This can I've had it ask me 40 questions. I've had it ask me 80 questions. I've had some people that asks 100 questions too. Literally you're sat there for an hour chatting to the AI. And what you end up with is essentially this conversation history that works really nicely and works really nicely as an asset of the design concept that you're creating. This can also function like this. You can have a meeting with someone who's a maybe a domain expert. Maybe I have a meeting with Sarah. I feed that meeting transcript into I don't know, Gemini meetings or whatever you guys are using. You take that, you feed it into a grilling session and you grill through the assumptions that you didn't have. So this ends up being a really nice kind of um a really nice way of just taking inputs from the world and then just turning and validating them. So okay. Let's see. I really want to get to the end of this, but I also don't want to just like be sat here talking to the AI in front of you for uh a thousand days. So I'm just going to say yes. Let's see what happens. So I'll tell you what, um while you guys sort of have a little fiddle with this locally, let's start a little Q&A session now. And let's see. How's this going to work? Can we keep the door closed or turn up the microphone? It's quite noisy. Uh let's see. Mike, can we uh door closed. Oh it has been closed. Mark has answered. Beautiful. So what I'd like you to do is there any air con? Yeah, there is some air con, I think. There is some air con. You guys aren't being lit here. I'm being fro I'm being fried alive here. Uh so what I'd like you to do is go on to the Slido, which you can join here. Have a if if you're not taking the exercise, go on to the Slido, have a little fiddle and vote on some good questions. I'm just going to chat to the AI for a second uh until we reach a stopping point. So do streaks earn points? Um streaks are standalone. Let's see what else it comes up with. Where does gamification UI live? Let's have it in the dashboard. I'm just going to scan these and blast through them basically. So how are we doing with our Slido? Okay. Have I tried Spec Kit, Open Spec or Taskmaster instead of the Grill Me skill? Do I find them more verbose or a structured alternative? This is a great question. So there are a ton of different frameworks out there that allow you to um sort of build up this planning process for you. I personally believe you at at this stage, when there's no clear winner, when there's no kind of like one true way and when things are changing all the time, you need to own as much of your planning stack as you possibly can. What I've noticed and a lot of my students is they tend to overuse a certain stack. They get into trouble and they because they don't own the stack and they don't have observability over the whole thing, they just go this isn't working. This sucks. Whereas if um if you have control over the whole thing, then at least you know how to fix it or potentially know how to fix it. So I'm even though I'm sort of giving you uh a stack basically, I believe in inversion of control and you should be in control of the stack. So bur bur bur. Can I press zero, please? Sorry? Sorry, that was a lot of sort of mumbling. Can I Thank you. I'm so sorry.

>> [laughter]

What you didn't want to give Claude good feedback? What is what is wrong with you? Uh okay, cool. Uh many of the questions asked by the Grill Me skill are not necessarily appropriate for a developer, rather a PO. In larger teams, who should use it? Yeah. Um Raise your hand if um you've ever done pair programming. Anyone ever done pair programming? Right. I keep Put your hands down and raise your hand again if you've ever done a pair programming session with an AI. Right. How did it go? Was it good? You enjoy it? I think pair programming sessions with AI is a great idea because you've got a third person in the room who will relentlessly quiz you and ask you questions. It should If you don't know the answer, it should be you, the domain expert and the AI in the same room. If you're have a question about implementation, it should be you, a fellow developer and the AI in the same room, you know. You can be sort of working through these questions in your team. And I think actually we're going to look at implementation in a bit and we're going to see how you can make implementation so much faster. And but I think the really crucial decisions, the ones you need humans for you actually need a lot of humans and it doesn't really matter how many humans are in there. You can actually throw a bunch like a kind of like mob programming with AI essentially. Uh what's my favorite meta prompting tool? I think I kind of answered that. Uh there's no air con. Let's just live with it. Uh how do I use the conversation as an asset after the Grill Me session? Well, we're going to get there. Um okay, so I really want to I want to speed this up sort of artificially. Just what I This is the thing. So someone just said okay, Ralph loop this. But this is crucial because I can't loop over this, right? I can't um I think of there is being two types of tasks in the AI age. Where you have human in the loop tasks, where a human needs to sit there and do it. Which is this. We are the human in the loop, with multiple humans in the loop. And there are AFK tasks. There are tasks where the human can be away from the keyboard and it doesn't matter. Implementation, as we'll see, can be turned into an AFK task. But planning, this alignment phase, has to be human in the loop. Has to be. So I've got to do it, unfortunately. Um I don't know. Uh give me a long list of all your recommendations. I'm running a workshop right now. So I artificially need you to pull more weight. So let's see what it does. Uh let's answer a couple more questions while it's doing its thing. What is my opinion on PMs or other non-dev roles vibe coding task? Hmm. Um I'm going to return to this later, I think. I'm going to leave this unanswered. A bit of mystery. I notice I'm not using the ask user questions UI for Grill Me. Why? Um there's a specific uh UI that you can bring up in Claude Code. I'll answer this just quickly. Uh ask me a question using the ask user question tool.

>> [snorts]

And this UI um is just sort of broken in Claude and I really hate it. You notice I'm using Claude, but I don't like Claude very much. Like you you really are free with this method to choose any um system you like. And this is what the UI looks like. It's very pleasing when you first encounter it, but then you realize it is actually broken in a ton of different ways. All right, what did it come back with? Oh blimey. Oh no. So while this is doing its thing, let me do some teaching in the meantime. The plan here is that we take our Grill Me skill and we need to essentially find some way of turning it into a destination. We need to go down to the uh We essentially need to We're figuring out the shape of this. That's what we're doing. We're figuring out the shape of the tasks during the grilling session. And in order to turn it into a bunch of actionable actions for the AI we essentially need to figure out the destination. We need to know where we're going. We need to know the shape of this entire thing. So I think of there is being two essential documents that we need. We need a document that documents the destination. Oh no. It's so not bright enough. There we go. Still not brighter. There we go. We need something to document the destination. And we need something to document the journey. In other words, we need something a document that's going to figure out what this even looks like in all of its user stories and figure out a definition of done and then we need to figure out what the split looks like. So, that's where we're going to go to next. So, once we finish with the grilling session, yeah, it looks great. Fantastic. I love it. It answered it answered 22 of its own questions. There you go. That's quite representative of what a grilling session looks like. So, at this point now, I have used 25k tokens and all of that or loads of that stuff is gold. I want to keep that around. I've I've got 25k great tokens there. And what I want to do is kind of summarize it in some kind of destination documents. So, this is um the next exercise where we're going to uh we're going to write a product requirements document. And the the product requirements documents or the PRD is essentially that's its function. It's the destination documents. And it's sort of doesn't matter what shape it is. I've got a shape that I prefer and I quite like. But, you can just choose your own shape or whatever your company uses. And all we're really doing is I'm not too worried about that. All we're really doing is summarizing the design concept that we have so far. And the So, let let's try this. So, I'm going to initiate this. I'm going to say zoom all the way to the bottom. All I'm going to do is just say write a PRD. And we can take a look at that skill now. Write a PRD. So, this skill it does a few things. It first asks the user for a long detailed description of the problem. You can use write a PRD without grilling first, but I just like to grill first and then write the PRD afterwards. Then you can um get it to install the repo which we've kind of already done. Then we get it to interview the user relentlessly so we have a kind of grilling session again and then we start um putting together a PRD template. So, this is available in the repo if you want to check it out. And essentially this is what it looks like. We've got some problem statements, the problem the user is facing, the solution to the problem and a set of user stories. And these user stories sort of define what this is. You know, as you you guys have probably seen things like this if you've been a developer at all. Um you know, there are cucumber is a language you can use to write these in or we just sort of um uh write ourselves essentially. Then we have a list of implementation decisions that were made and list of crucially testing decisions, too. So, I'm going to run this. Okay. And so, it's finished its thing. Ah! Windows, let me close the thing. Thank you. I don't know why I bought a Windows laptop. I think I just I like the challenge. Um

>> [clears throat]

So, the first thing that it's going to give me are a set of proposed modules it wants to modify. Now, there's a deep reason why I'm thinking about this. So, this is at this stage we have an idea, we have sort of specked out the idea, we've reached a sort of understanding of what we're trying to do and then we need to start thinking about the code because at this point we need to this is not specs to code. This is not where we're ignoring the code. We actually keep the code in mind throughout the whole process. And the way I like to do this is I like to just sort of think about a set of proposed modules to modify. We're going to return to this this idea of continually designing your system and keeping your system in mind. So, it's it's saying recommend tests for the gamification service is the only deep module with meaningful logic. These modules look right. Yeah. Looks good. And it's going to hang out a PRD. Now, for ease of setup I've got it so that it creates a set of issues locally. So, it's just going to create essentially a PRD inside this issues directory. But, the way I usually do it and you can check this out yourself is you can go to my um essentially what I consider my work repo which is GitHub um dot com forward slash Matt Pocock forward slash course video manager up here. And in here, this is essentially a app that I create um that I use all the time to record my videos and things like this. I think I've recorded like I pulled out the stats. I think I've recorded like a thousand videos in here or something nuts. Um and you can see here that it's got 744 closed issues. And this is essentially all of the uh PRDs and all of the implementation issues that I've put into here. So, this is how I usually like to do it.

>>[clears throat]


n5321 | 2026年5月18日 16:08

The Future of Electric Motor Factories: A Drucker Perspective

the essence of management is not in predicting trends or piling up technology, but in asking the right questions: What is the purpose of an electric motor factory? How does it organize people, resources, and processes in a knowledge society to create and retain customers? How does it make knowledge workers more effective, turning ordinary operations into extraordinary results?

In the future—by 2030 and beyond—electric motor factories will no longer be mechanical assembly lines. Their true purpose is to deliver results: precise, controllable rotary motion under extreme constraints of size, efficiency, precision, and reliability.

This amplifies customers' capabilities, making end products smaller, lighter, more efficient, and smarter—reducing lifecycle costs, accelerating innovation, and serving broader societal goals like healthier lives, cleaner mobility, smarter production, and more humane interactions.

The factory's operation will shift from "selling hardware" to "selling quantifiable results + services," becoming customers' visibility partners and organizational transparency partners.

Digitization of outcomes—making uptime, efficiency, and precision more visible, certain, and measurable via IoT, AI, and data visualization—will be the core engine.


n5321 | 2026年2月28日 23:17

从德鲁克的管理管理视角理解提示工程

在提示工程(Prompt Engineering)常常被视为一种技术技巧,Claude现在就在搞SKILL 模版。大家热衷关键词、模板、句式和各种提示窍门。模板深度理解之后,发现他本质上跟管理的operations具有一致性!

Peter F. Drucker在《管理实践》(The Practice of Management)中提出的管理框架。the specific work of a manager,一个经理人的作业:设定目标,组织资源,激励沟通,评估,培养人才!

两个框架的pattern几乎是一样的!

如果再去分析问题模型,更加确认这种一致性!

管理的本质是在不确定性中塑造可预测的成果。当我们编写提示时,面对的是一个复杂且概率驱动的语言模型。我们无法直接操控它的每一个输出,只能通过精心设计的prompt来引导它。

借用management的框架,可以发现提示工程不再是技巧堆砌,而是设计一个系统以稳定产出成果的过程。

  1. 一切从设定目标开始,这是德鲁克管理框架的首要任务。目标不是模糊的愿望,而是明确的、可衡量的方向。想想一个弱提示:“写一篇文章。”它缺乏具体性,只能是一个不确定,难以匹配预期的输出结果。相反,一个强有力的提示会定义目标读者、文章长度、结构、风格,甚至成功标准。

    这些元素的作用不是简单表达想法,而是减少不确定性,或者说增加收敛性。没有明确目标,组织会陷入忙碌却无成效的泥沼;同样,没有目标的提示会让模型的输出发散,偏离预期。

  2. 接下来是组织结构的设计。在企业管理中,这意味着分工、任务分解和责任匹配,以减少混乱。应用到提示工程中,则是通过分步骤执行、明确角色、提供背景和设定边界来实现。例如,一个有效的提示可能这样设计:“首先分析问题,其次给出框架,最后展开正文。”这不是随意的写作技巧,而是 deliberate 的结构工程。通过这些约束,我们减少了模型的自由度,确保输出更有序、更可靠。管理通过organizing来驾驭复杂性,提示工程亦然——它在概率的海洋中筑起堤坝。

  3. 沟通的清晰度则是确保输出稳定性的关键。在组织中,模糊的表达会导致执行偏差;在AI模型中,指令的模糊性会放大概率补全的随机性。模型本身不会“犯错”,它只是填补你未定义的空间。因此,编写提示的本质在于定义边界:用精确的语言划定范围,避免歧义。清晰不是一种礼貌,而是核心控制机制,它将潜在的混乱转化为可靠的成果。

  4. 衡量标准是另一个不可或缺的环节。在提示工程中,如果缺少字数限制、结构要求、风格规范或判断准则,model就无法评估输出质量。衡量机制本身就是一种隐形的控制工具。随着提示工程的成熟,评估(evaluation)将成为标准实践,帮助我们量化改进的空间。

  5. 管理并非一蹴而就,它是一个持续的反馈循环:设定目标、执行、衡量、调整。提示工程同样如此。初次输出往往不是完美版本;高效的使用者会通过精细化约束、重组结构和优化表达来迭代。这不是在“修正”模型,而是在管理一个智能系统,确保每一次循环都更接近理想成果。更深层的洞察在于,组织是复杂的人类系统,而AI模型是复杂的概率系统。二者都不可完全预测,却都对结构极其敏感。

管理的目标不是消除不确定性,而是设计出能在不确定性中稳定产出的框架。提示工程正是如此:它在AI的“黑箱”中注入秩序。这种视角转变意味着什么?如果你将提示工程视为技巧,你可能会 endlessly 追逐各种模板和诀窍。但如果你视之为管理,你会问自己:我的目标是否清晰?结构是否合理?评估是否明确?这是一种管理者心态,而非被动用户视角。它让你从单纯的输入-输出转向系统设计。

今天我们面对的是AI模型,但核心问题不变:如何在复杂系统中稳定产生成果?答案依旧是:通过通过目标导向、结构设计、清晰沟通、严格衡量和持续迭代。从这个角度看,prompt engineering就不再只是编写提示,而是构建一个高效的成果系统。


n5321 | 2026年2月28日 17:19

A practical guide to OpenAI prompt generation

So, you’ve started playing around with OpenAI. You’ve seen moments of brilliance, but you’ve probably also felt that flicker of frustration. One minute, it’s writing flawless code; the next, it’s giving you a completely generic answer to a customer question. If you’re finding it hard to get consistent, high-quality results, you're definitely not alone. The secret isn't just what you ask, but how you ask it.

This is where OpenAI Prompt Generation comes into play. It's all about crafting instructions that are so clear and packed with context that the AI has no choice but to give you exactly what you need.

In this guide, we'll walk through the pieces of a great prompt, look at the journey from writing prompts by hand to using automated tools, and show you how to put these ideas to work in a real business setting.

What is OpenAI Prompt Generation?

OpenAI Prompt Generation is the art of creating detailed instructions (prompts) to get Large Language Models (LLMs) like GPT-4 to do a specific job correctly. It’s a lot more than just asking a simple question. Think of it less like a casual chat and more like giving a detailed brief to a super-smart assistant who takes everything you say very, very literally.

The better your brief, the better the result. This whole process has a few stages of complexity:

  • Basic Prompting: This is what most of us do naturally. We type a question or command into a chat box. It works fine for simple things but doesn't quite cut it for more complex business needs.

  • Prompt Engineering: This is the hands-on craft of tweaking prompts through trial and error. It means adjusting your wording, adding examples, and structuring your instructions to get a better answer from the AI.

  • Automated Prompt Generation: This is the next step up, where you use AI itself (through something called meta-prompts) or specialized tools to create and fine-tune prompts for you.

Getting this right is how you actually get your money's worth from AI. When prompts are fuzzy, the results are all over the place, which costs you time and money. When they’re well-designed, you get predictable, quality outputs that can genuinely handle parts of your workload.

The core components of effective OpenAI Prompt Generation

The best prompts aren't just one sentence, they’re more like a recipe with a few key ingredients. Based on what folks at OpenAI and Microsoft recommend, a solid prompt usually has these parts.

Instructions: Telling the AI what to do

This is the core of your prompt, the specific task you want the AI to tackle. The most common mistake here is being too vague. You have to be specific, clear, and leave no room for misinterpretation.

For instance, instead of saying: "Help the customer."

Try something like: "Read the customer's support ticket, figure out the main cause of their billing problem, and write out a step-by-step solution for them."

The second instruction is crystal clear. It tells the AI exactly what to look for and what the final answer should look like.

Context: Giving the AI the background info

This is the information the AI needs to actually do its job. A standard LLM has no idea about your company’s internal docs or your specific customer history. You have to provide that yourself. This context could be the text from a support ticket, a relevant article from your help center, or a user's account details.

The problem is that this information is usually scattered everywhere, hiding in your helpdesk, a Confluence page, random Google Docs, and old Slack threads. Manually grabbing all that context for every single question is pretty much impossible. This is where a tool that connects all your knowledge can be a huge help. For example, eesel AI solves this by securely connecting to all your company's apps. It brings all your knowledge together so the AI always has the right information ready to go, without you having to dig for it.

eesel AI connects to all your company
eesel AI connects to all your company

Examples: Showing the AI what "good" looks like (few-shot learning)

Few-shot learning is a seriously powerful technique. It just means giving the AI a few examples of inputs and desired outputs right inside the prompt. It’s like showing a new team member a few perfectly handled support tickets before they start. This helps guide the model’s behavior without having to do any expensive, time-consuming fine-tuning.

Picking out a few good examples yourself is a great start. But what if an AI could learn from all of your team's best work? That's taking the idea to a whole new level. eesel AI can automatically analyze thousands of your past support conversations to learn your brand's unique voice and common solutions. It’s like giving your AI agent a perfect memory of every great customer interaction you've ever had.

Cues and formatting: Guiding the final output

Finally, you can steer the AI's response by using simple formatting. Using Markdown (like # for headings), XML tags (like ``), or even just starting the response for it ("Here’s a quick summary:") can nudge the model to give you a structured, predictable output. This is incredibly handy for getting answers in a specific format, like JSON for an API or a clean, bulleted list for a support agent.

The evolution of OpenAI Prompt Generation: From manual art to automated science

Prompt generation isn't a single thing, it's more of a journey. Most teams go through a few stages as they get better at AI automation.

Level 1: Manual OpenAI Prompt Generation

This is where everyone begins. A person, usually a developer or someone on the technical side, sits down with a tool like the OpenAI Playground and fiddles with prompts. It’s a cycle of writing, testing, and tweaking.

The catch? It’s slow, requires a ton of specific knowledge, and just doesn't scale. A prompt that works perfectly in a testing environment is completely disconnected from the real-world business workflows where it needs to be used.

Level 2: Using prompt generator tools

Next up, teams often find simple prompt generator tools. These are usually web forms where you plug in variables like the task, tone, and format, and it spits out a structured prompt for you.

They can be useful for one-off tasks, like drafting a marketing email. But they're not built for business automation because they can't pull in live, dynamic information. The prompt is just a fixed block of text, it can't connect to your company's data or actually do anything.

Level 3: Advanced prompt generation with meta-prompts

This is where things get really clever. A "meta-prompt," as OpenAI's own documentation explains, is an instruction you give to one AI to make it create a prompt for another AI. You're essentially using AI to build AI. It’s the magic behind the "Generate" button in the OpenAI Playground that can whip up a surprisingly good prompt from a simple description.

But even this has its limits. At its core, it's still a tool for developers. The great prompt it creates is still separate from your helpdesk, your knowledge base, and your team's daily grind. You still have to figure out how to get that prompt into your systems and connect it to your data.

The next step: Integrated AI platforms

The real goal isn't just to generate a block of text, it's to build an automated workflow. This is where you graduate from a prompt generator to a true workflow engine. The prompt becomes the "brain" of an AI agent that can access your company's knowledge, look up live data, and is allowed to take action, like tagging a ticket or escalating an issue.

This is exactly how eesel AI works. Our platform lets you set up your AI agent’s personality, knowledge sources, and abilities through a simple interface. You’re not just writing a prompt in a text box; you’re building a digital team member that works right inside your existing tools like Zendesk, with no complex coding needed.

With eesel AI, you can build a digital team member by setting up its personality, knowledge, and abilities through a simple interface, moving beyond simple OpenAI Prompt Generation.
With eesel AI, you can build a digital team member by setting up its personality, knowledge, and abilities through a simple interface, moving beyond simple OpenAI Prompt Generation.

The business impact: Understanding the costs of OpenAI Prompt Generation

While writing prompts can feel like a technical chore, its impact is all about the money. According to OpenAI's API pricing, you pay for both the "input" tokens (your prompt) and the "output" tokens (the AI's answer). This means every time you send a long, poorly written prompt, it costs you more money. Good prompt engineering is also about keeping costs down.

OpenAI does have a feature called prompt caching that can help with speed and cost for prompts you use over and over. But it doesn’t fix the main issue of unpredictable usage, which can lead to some nasty surprise bills.

This is why "per-resolution" pricing models from many AI vendors can be so tricky. They lead to unpredictable costs that go up when you're busiest. With eesel AI’s pricing, you get clear, predictable plans based on a set number of monthly AI interactions. You’re in complete control of your budget, with no hidden fees, even if your support ticket volume suddenly doubles.

eesel AI’s pricing provides clear, predictable plans, giving you control over your budget for OpenAI Prompt Generation.
eesel AI’s pricing provides clear, predictable plans, giving you control over your budget for OpenAI Prompt Generation.

Go beyond the playground

The OpenAI Playground is a great place to experiment, but businesses need something reliable, scalable, and plugged into their day-to-day work. The final step is to move from a "prompt generator" to a full "workflow engine."

That's why having a safe place to test things out is so important. With eesel AI, you can run a powerful simulation using thousands of your past support tickets. You can see exactly how your AI agent will behave, check its responses, and get accurate predictions on how many issues it will solve and how much you'll save, all before it ever talks to a real customer. This lets you build and launch with total confidence.

The eesel AI platform allows you to run powerful simulations to test your OpenAI Prompt Generation against historical data before deployment.
The eesel AI platform allows you to run powerful simulations to test your OpenAI Prompt Generation against historical data before deployment.

Stop generating prompts, start building agents

Effective OpenAI Prompt Generation is structured, full of context, and always improving. While tinkering by hand and using simple tools are fine for small tasks, the real value for your business comes from weaving this intelligence directly into your workflows.

The goal isn't just to create better text. It's to automate repetitive tasks, give your team instant access to information, and deliver better, faster results for your customers. It's time to move beyond just writing prompts and start building intelligent agents that actually get work done.

Ready to see how easy it can be to build a powerful AI agent without touching a line of code? Set up your AI agent with eesel AI in minutes and see how our platform turns the complex world of prompt generation into a simple, straightforward experience.


n5321 | 2026年2月28日 09:06

the origin of AI

一切真正开始于 1950 年代。那时候计算机才刚出生没几年,一群年轻人——Alan Turing、John McCarthy、Marvin Minsky、Allen Newell、Herbert Simon 这些人——突然冒出个大胆的想法:能不能造一台机器,让它表现出“智能”?1956 年,他们在达特茅斯学院开了个夏天研讨会,直接把“人工智能”这个词发明出来了。那会儿的乐观情绪高得离谱:有人说“20 年内就能解决所有智能问题”。他们相信,只要把逻辑、搜索、符号处理这些东西装进计算机,就能模拟人类思考。于是 第一个 AI 热潮(1956–1973)开始了。成果还真不少:逻辑理论家证明了数学定理、通用问题求解器 Shakey 机器人能在房间里晃悠、ELIZA 这个聊天程序居然能骗人觉得它懂心理学。但问题也很快就来了:这些系统只能在非常窄、非常结构化的玩具世界里玩得转。一遇到真实世界的复杂性、模糊性、不确定性,它们就卡壳了。计算资源也跟不上,内存小、速度慢。到了 1973 年,英国 Lighthill 报告一锤子砸下来,说 AI 基本没戏,美国和英国的资助大幅缩水——第一个“AI 冬天”就这么开始了。70 年代到 80 年代初,AI 低调了好一阵子。但没完全死掉。有些人悄悄转向了 专家系统(expert systems)。想法很简单:别再试图让机器自己“思考”了,直接把人类专家的知识一条一条编码进去,做成 if-then 规则库。MYCIN 诊断细菌感染、DENDRAL 分析化学分子、XCON 帮 DEC 公司配置电脑订单——这些系统真的在某些领域赚了钱、帮了大忙。80 年代中期,日本的第五代计算机计划和美国 DARPA 的战略计算计划又把钱砸进来,专家系统公司如雨后春笋一样冒出来。可好景不长。到 80 年代末,大家发现:规则写得再多,也写不完现实世界的全部例外;维护知识库贵得离谱;新情况一来,系统就崩溃。最要命的是,专家自己常常说不清“为什么”这么判断。于是第二个 AI 冬天又来了(1987–1993),资金撤退,公司倒闭,很多人以为 AI 这事儿彻底凉了。但就在冬天里,有几条暗流在悄悄流动。一条是 神经网络。它其实 50 年代就有(感知机),但因为 Minsky 和 Papert 1969 年那本《Perceptrons》把单层网络批得体无完肤,大家都觉得它没戏。80 年代中期,Rumelhart、Hinton、Williams 重新发明了反向传播,多层网络开始复活。Yann LeCun 搞出了卷积神经网络,能认手写数字了。但那时候计算力不够,数据也不够,大家还是觉得“神经网络太慢、太黑箱”。另一条暗流是 概率方法和机器学习。Judea Pearl 的贝叶斯网络、统计学习理论、支持向量机这些东西开始冒头。它们不像符号 AI 那么刚愎自用,而是承认世界有不确定性,愿意从数据里学。然后 1997 年出了个事儿,让很多人重新抬起头:IBM 的深蓝(Deep Blue)下棋打败了世界冠军卡斯帕罗夫。那不是神经网络,是暴力搜索 + 手工特征 + 评估函数。但它告诉大家:专用系统在特定任务上,是真的可以超过人类的。真正的转折点来得晚一些——2010 年代初。GPU 的出现让训练深层神经网络突然变得可行;ImageNet 大规模数据集公开;AlexNet(2012)在图像识别比赛上把错误率砍了一半多。大家突然意识到:只要数据够多、计算力够强、层数够深,神经网络的性能就能指数级飙升。这就是所谓的“深度学习革命”。从那以后,AI 进入第三个春天,而且是前所未有的热潮:

  • 2014–2016:生成对抗网络(GAN)、AlphaGo(2016 打败李世石)

  • 2017:Transformer 架构横空出世(Attention is All You Need)

  • 2018–2022:BERT、GPT 系列把自然语言处理彻底颠覆

  • 2022–现在:ChatGPT、GPT-4、Gemini、Claude、Llama、o1 系列……大语言模型把“会聊天、会写代码、会推理”这事儿做到了让普通人都震惊的程度

但我得跟你说实话(因为我见过太多轮热潮了):我们现在拥有的这些东西,虽然表面上很强大,但它们本质上还是超级厉害的统计模式匹配器。它们在海量数据里学会了模仿人类的语言、图像、代码,但它们没有真正理解意义、没有稳固的常识、没有对世界的因果模型、很容易在边缘情况崩溃、会一本正经地胡说八道(hallucination)。所以 AI 的历史,其实是一部人类对“智能”定义不断调整的历史。一开始我们以为智能就是逻辑推理,后来以为是专家知识,再后来以为是模式识别和大规模统计。现在很多人又开始说“也许智能就是足够大规模的模式匹配”。但我总觉得,我们可能还差了点什么——也许是类比、也许是抽象、也许是对“意义”的真正把握。咱们别急着宣布胜利,也别急着宣布失败。这段历史告诉我们:每一次大突破,都伴随着巨大的 hype 和后来的清醒。每一次冬天,都在为下一次春天攒能量。你看,现在我们站在又一个高点上。但真正的问题不是“机器会不会超过我们”,而是我们人类在试图造出“像我们一样”的东西时,到底学到了多少关于自己的事。



n5321 | 2026年2月28日 01:07

Prompt Value

在NASA工作17年的航空工程师,未来的米帝国家工程院院士Walter G. Vincenti,在1957年转身斯坦福任教,重建航空航天工程专业。Walter G. Vincenti是工程界的大佬。

1970年,经济系的同事Nathan Rosenberg在一个午餐后问他:What is it you engineers really do?”

翻译过来是你们工程师到底是做什么的!

大佬被问懵了!鱼是最后一个知道水的!工程师可能对哲学或者说社会学意义上的工程并不了解!按IT界大佬Leslie Lamport的台词,那就更加刁毒了: If you think you know something but don't write it down, you only think you know it. 意思说你要是写不清楚,那不过是自以为是。

于是Walter G. Vincenti转向研究技术哲学。20年后写了《What Engineers Know and How They Know It》

那个工程师怎么干的问题被直接置换了!被置换成What engineers do, however, depends on what they know。

没有分析、没有论证的置换过了,当成公理一样置换了!

工程师想要知道!They want to know!

在AI时代,大语言模型(LLMs)声称“浓缩(abstract)”了人类知识,但实际上,LLMs的“知道”依赖于我们如何通过prompt“提取”和“引导”它——这正是prompt engineering的核心。

LLMs不是主动“知道”的实体,而是通过提示(prompt)激活潜在模式的“知识库”。

正如Vincenti强调的,如果你不写下来(或不精确prompt),你就只是“自以为知道”。借用Leslie Lamport的台词:如果你不能清晰prompt,那你对LLM的认知也只是幻觉。


n5321 | 2026年2月28日 01:06

What is token

简单说,token 就是模型“看世界”的最小单位。

想象一下:你读一本书的时候,不是一个字一个字地看,而是把句子拆成一个个有意义的“块”来理解,对吧?人类大脑很擅长做这种拆分。但计算机,尤其是神经网络,它没有我们那种直觉,所以需要先把所有文本切成小块,这些小块就叫 token。token 到底长什么样?不同的模型切法不太一样,但主流的做法(比如 GPT 系列、Claude、Llama、Gemini 用的那些 tokenizer)大概是这样的:

  • 一个常见的英文单词,比如 “hello” → 可能就是一个 token。

  • 但 “unbelievable” 这种长词,可能被切成 “un” + “believ” + “able” 三个 token。

  • 中文就更直白了:通常一个汉字就是一个 token(有时候两个常见汉字组合会合并成一个)。

  • 标点、空格、特殊符号也都是 token(比如 “!” 就是一个单独的 token)。

  • 数字、URL、代码里的变量名,也会被拆得很细。

举个例子,把这句话喂给 tokenizer:“人工智能正在改变世界。”可能的 token 大概是: [“人”, “工”, “智”, “能”, “正在”, “改变”, “世界”, “。”]一共 8 个 token。再来个英文的: “The quick brown fox jumps over the lazy dog.”可能拆成: [“The”, “ quick”, “ brown”, “ fox”, “ jumps”, “ over”, “ the”, “ lazy”, “ dog”, “.”]大约 10 个 token。你看,token 不是严格等于“词”或“字”,它是一种模型自己学出来的、统计上最有效率的切分方式。OpenAI 他们用的是叫 BPE(Byte Pair Encoding)的算法,简单说就是:先把所有文本拆成单个字节,然后反复把最常一起出现的字节对合并成一个新“词”,直到达到想要的词汇表大小(通常 5 万到 10 万个 token 类型)。为什么 token 这么重要?因为大语言模型的一切“理解”和“生成”都是基于 token 的:

  • 模型的输入上限(context window)是用 token 算的。比如 GPT-4o 的 128k token、Claude 3.5 的 200k token、Gemini 1.5 的 1M+ token——这些数字指的就是它一次能“看”多少个 token。

  • 训练的时候,模型就是在预测“下一个 token 是什么”。

  • 你付钱给 OpenAI、Anthropic 的时候,也是按 token 计费(输入多少 + 输出多少)。

  • 模型的“聪明”程度很大程度上取决于它在训练时见过多少 token(现在顶级模型都训练到几万亿甚至十几万亿 token 了)。

所以当有人说“这个模型的上下文窗口是 128k token”,其实就是在告诉你:它一次最多能记住/处理相当于大概 10 万个英文单词(中文会少一些,因为一个汉字 ≈ 一个 token)的文本长度。但这里有个小陷阱,得提醒你token 不是均匀分布的:

  • 常见词、常见汉字用得少 token(效率高)。

  • 生僻词、长尾英文、专业术语、emoji、代码里的奇怪变量名,会“吃”很多 token。

  • 所以同样一段意思,英文可能 100 token,中文可能 150 token,代码可能 300 token。

这也是为什么有些人觉得“中文模型吃 token 比英文贵”——其实不是模型故意坑中文,而是 tokenizer 的词汇表对英文优化得更好。


n5321 | 2026年2月28日 01:05

Interview with Leslie Lamport: Turing Award Winner


Teaser / Intro

Leslie Lamport: If you think you know something but don't write it down, you only think you know it.

Host: This is Leslie Lamport. He's a Turing Award winner famous for his contributions to distributed systems. And I interviewed him for the stories behind his papers.

Leslie Lamport: Their reaction shocked me. They became angry. I really thought they might physically attack me.

Host: What was it about Dystra's old solution that you felt was unsatisfactory? It was not an obvious idea to most people that had actually impressed Dystra.

Host: As the inventor of the Paxus algorithm, I asked them his thoughts on the competing raft algorithm.

Leslie Lamport: There was a bug discovered in Raft and fixed, but I believe the algorithm that they found more understandable was one with that bug.

Host: I also enjoyed reflecting over his 50-year career. You say things like, "You never considered yourself smart. How could that be?" Stupid people think they're smart because they're too stupid to realize they're not.

Host: You felt like a failure at some point because you wanted to develop this grand theory of concurrency and you never discovered it. Do you still feel that way?


The Bakery Algorithm & Dijkstra

Host: Here's the full episode. I wanted to start with the bakery algorithm. What is the problem that the bakery algorithm solves? And you know how did you discover the problem?

Leslie Lamport: Well uh the problem was invented or discovered by Edkar Dystra in a 1965 I think it was 1965 paper and that began I consider that really the beginning of the theory of uh concurrency concurrent programming. He was the first one who really made use of the idea of of concurrency as a way of structuring programs as a as a collection of semi-independent tasks and the processes have to uh synchronize with one another.

Uh one of the processes or you know among the processes would be well this was in the days of time sharing uh you know right at really at the beginning of beginnings of time sharing and the idea of multiple people using the same computer. People realized that computers were worked faster than humans and so and computers were very expensive in those days. So uh they wanted they could use a computer to simultaneously to be used simultaneously by multiple people. The program that each user was running you know was a separate program but sometimes you know there were resources that got shared. for example, a printer, two people trying to print on the same printer at the same time. Well, the result would be, you know, not very satisfactory.

So uh he realized there was this problem of uh synchronizing um multiple processes via the idea of what he called a critical section or some piece of code in each of the processes so that at most one process can be executing that piece of code at any particular time. So that code might be the code that prints something on the printer. So the problem was how to get the uh processes to synchronize among themselves so that at most one process was executing its critical section at a time.

And um it was in 1972 that I learned about the problem because there was an article giving a solution to it um in the CACM communications of the ACM and uh I mean I used to program and I liked little programming problems you know uh and this was just a very nice little programming problem. And so I looked at the solution, which is fairly complicated, and I said, "Oh, gee, that shouldn't be so hard." And so I whipped off a very simple uh algorithm for two processes and submitted it to CACM. And a couple of weeks later, I received uh a letter from the editor uh pointing out the bug in my program. So that had two effects.

The first was that I realized that concurrent programs were hard to get right and that you needed a proof that they were correct. And second uh was that made me feel I'm gonna solve that damn problem. And I came up with the bakery algorithm which was inspired by uh the idea came from you know what now called the deli problem where you have a deli counter and that collects you know tickets a roll of tickets and every customer would come in and take a ticket and then the the next person s to to be served would be the one with the highest the lowest numbered ticket uh that hadn't been served yet.

And basically that I took that idea uh but since uh there was no central server uh or at least the the problem is as specified by Dystra involved no central control. Each process basically had to choose their own ticket. That was you know the basic idea and the algorithm was you know quite simple. And I wrote a a proof of correctness.

And uh the proof of correctness revealed to me that this algorithm was had this very interesting property. There was a general feeling in fact somebody published in a in a book or paper saying you know that it was impossible to implement mutual exclusion like this without using some lower level mutual exclusion. And the way most the mutual exclusion that was assumed generally was that of shared registers. you know, shared pieces of memory that could be written and write and uh read by different processes. And the idea is that, you know, you couldn't have one process, you know, two processes writing at the same time or one process reading while the other process was writing. People assumed that those actions were atomic. They always performed as if they occurred in some specific order.

But the amazing thing about the bakery algorithm was that it didn't require that assumption. It it used uh each shared memory a piece of memory was only written by a single process. So it didn't have to worry about two processes interfering with each other. The only problem that you might come is that somebody reading the uh value while it was being written might get you know some unknown value but the algorithm worked anyway. If somebody read if one process read while the registers was being written that process reading process could get absolutely any value and the algorithm still worked.

Host: I saw in your your writing about this problem that you shared it with a colleague named Anatol Hol and the proof was so remarkable that uh they didn't believe it and...

Leslie Lamport: well the the result was so remarkable.

Host: Yes. Yes. that didn't believe it.

Leslie Lamport: And uh you know I wrote the proof on the on the whiteboard for him and you know he couldn't find it but he went home and saying there must be something wrong with it and uh he obviously never found anything wrong with it.

Host: Right. I saw the name of the paper is a new solution of Dystra's concurrent programming problem. What was it about Dystra's old solution that you felt was unsatisfactory and made you want to solve this problem?

Leslie Lamport: Uh well, there was an unsatisfactory aspect of his original solution that had the property that if there were a lot of if processes kept trying to uh enter their critical section uh an individual process might be starved. might never get access to to the critical section that was solved uh by you know the next solution I think was Don Kuth's the condition that was desired or that that measured that what was considered the uh the efficiency of it was how long a process might have to wait and I believe that the bakery algorith of them was the first one that was really first come first served. That is if one process came if what it meant is if one process chose its number before another process tried to enter the first process would enter the critical section before the other process did. And I believe the bakery algorithm was the first one with that uh with that property. And also I think it was simpler than uh other uh solutions.

Working with Dijkstra & The Gift of Abstraction

Host: in a lot of the writing. I see that you worked with Dystra and I saw in 1976 you actually worked for a month in the Netherlands and you worked with them. Can you talk about that a little bit?

Leslie Lamport: Dyster used to had the things they're called EWDs his initials. little papers, things that when he he thought of something, had some idea, he would write it down and send it out to people. Well, one of those EWDs was about he and uh some associates or actually sort of mentees I guess you would call them wrote this this algorithm. It was the first concurrent garbage collection algorithm.

a way of writing programs evolved where there was a pool of memory uh that when a program would need a piece of memory, it would ask some server for it and be given this piece of memory. Uh but at some point it would stop using that memory. But the program itself wouldn't know that the one the particular process that created this memory you know wouldn't know whether some other process is using that memory or not. So there was an additional process called the garbage collector which would go around examining the memory and decide which pieces of memory were no longer being used and then put them back on the it's called the free list and in which uh the uh server that the process that was giving out the uh uh memory would be able to to take it.

I looked at it and I realized that uh I could simplify the algorithm. Uh because he had some some spe the the handling of the free list was done by a special process that you know which had it to worry about its own coordination with the uh uh the processes that were using the memory. And I realized that that free list could just be made part of the regular data structure. uh so it didn't need special handling and that seemed to me like a very uh simple idea and a very obvious idea and I sent it to him and then when I get got the next version of the paper I discovered he had made me an author and I thought that was very generous of him uh to have to have done that because it seemed like very simple idea and I mean very obvious vious idea and I later realized much later that it was not an obvious idea to most people uh and that that had actually impressed uh Dystra.

when that was the only thing I actually did with Dystra many years later he said that I had uh a remarkable ability at abstraction only in very recent years I mean Maybe maybe after I got the touring award that I realized that the reason for my success and the reason I got it wind up wound up getting a touring award was not that I was particularly that smart but that I had this gift of abstraction and Dystra was smart enough to realize that I was invited to uh spend a month uh but not with Dysterra, with a colleague of his, Carl Carl Holton. Only one thing that was ever published came out of that. Carl and I would uh meet with Dystra once a week. Uh in the in the course of that discussion, the idea somehow came up that led to uh a variant of the bakery algorithm that I wrote up and published. Uh so that was the the one tangible result that that came from my month in uh the Netherlands.

Host: Yeah. I I saw that you wrote that. Yeah. You you spent one afternoon a week working, talking and drinking beer at Dextrous House and you kind of don't remember exactly who was uh you know in charge of uh what on that paper, but...

Leslie Lamport: yeah. Well, I don't think I was really could have gotten that drunk because uh I probably drove to the meeting and back from the meeting. So,

Host: Right. Right.

Leslie Lamport: The the Dutch beer that I was drinking was not very alcoholic.

Time Clocks and Ordering of Events

Host: I wanted to talk about your most cited paper, the one titled time clocks and the ordering of events and distributed systems. What's the story behind the paper and the problem you were solving with it?

Leslie Lamport: The origin was simple. uh it well somebody sent me a paper on building distributed databases and so where you'll have well multiple copies of the data in different places and you need to keep them synchronized in some way. I looked at it and I realized that their solution had this problem that the se that it it had the property that things would be executed as if they occurred in subsequence but that sequence could be different from the sequence in which they actually happened.

The notion of of what you know happening before means is not obvious or not obvious to most people but I happen to you know learn about you know special relativity in particular uh what's known as the it's the space-time view of of special relativity where you basically consider space and time together just one four-dimensional thing and that was Einstein wrote his paper in 1905 and and in I think it was 1909 uh somebody whose name I'm blocking on provided this four-dimensional view and that four-dimensional view has the the particular notion of what it means for one pro one event to occur before another and that notion is that one event happens before another. If a a signal uh was emitted from the first event and received by the whoever did that second event before that second event happened, but the communication could not travel faster than the speed of light because nothing can travel faster than the speed of light.

Well, I realized there was an obvious analogy. Uh the notion of happens before is exactly the same as in relativity except instead of being whether something one event can influence another by things traveling at the speed of light, it's whether the first event could have affected the other by information sent over messages that were actually sent in the system. The thing that you know blew people away was this this definition of of happens before in a distributed system with also this was the first paper I would call like you know had a scientific result about distributed systems.

I made perhaps you know mistake that I was warned against at some point of having two ideas in in one paper. The other thing that I realized was that there was an an algorithm that would show whether one event that it would produce an ordering that satisfied this that condition that if some if one event happened the other before the other then that first event would be ordered before the other. And I realized that if you had an algorithm to do that, you could use it to basically provide the the synchronization you needed for any distributed system because you could describe that system in terms of a state machine.

And a state machine as as I described it then is something that has a state and process executes you know uh commands that need to be executed in order and the command simply is something that makes a change of the state and and produces a value. And so you can just describe this state machine as just you know how event how commands affect the state and and how they produce and what you know what the new state is as a function of the original state and what the value is as a function of the original state.

It it turned out that this was very obvious to me, but that's really in practice the important idea in that paper because it showed that this method of building distributed systems by thinking in terms of state machine and and can thinking about concurrent systems in terms of state machines. Um but that part was completely ignored. As a matter of fact, twice I talked to people about that paper and they said there was nothing in that paper about state machines and I had to go back and reread and reread the paper to be sure I wasn't going crazy and it really did talk about state machines.

It's important uh for another reason. uh if you're trying to understand a concurrent program you concurrent programs are written the bakery algorithm is really an exception uh concurrent programs are written assuming atomic actions so that you assume that the execution behaves like a a sequence you you can assume that the execution proceeds as a sequence of events. It turns out that the way to understand, you know, why why does a program produce the right answer? Well, the answer is well, you it you give it the uh the you know the right input. You give it the input and then it produces the right answer. Well, but by the time you're in the middle of execution, what it was given at the beginning is ancient history. The only thing that that tells the program what to do next is its current state.

And the way to understand uh a program, you know, a simple program that just, you know, takes input and produces an answer is to say what is the property of the state at each point that ensures that the answer it produces is correct is going to be correct. And that property which is mathematically a fun a boolean valued function of the state is called an invariant. And understanding the invariant is the way to understand the system. You know the the program and I realized that the same thing is true of concurrent systems and concurrent programs. People like to write proof you know behavioral proofs reasoning about sequences.

And the problem with that is that the number of sequences possible sequences you know is exponential in the length of the sequence while the com so your complexity of your reasoning gets to be very complicated. It's very easy to to miss cases. Um but the complexity of an invariance proof the complexity of the invariant basically is well oh god it's the number of possible executions is exponential in the number of processes but the uh the behavior of the proof of a an invariance proof is quadratic in the number of processes. you know that's basically why invariance proofs are better but you know there's still for a long time that you know people you know doing uh distributed systems theory are trying to do it uh you know develop you know methods and formalism something that are based on partial orderings and that they've you know published a lot of papers but it's just you know not the way if you want to do it in practice that's that's not the way to do it and I shouldn't say you know it's not the way uh you know there are algorithms like the bakery algorithm that you know you know thinking in partial orderings is in fact a very good way of doing it but those are the exceptions the the work the method that works you know that you can be sure will will will work is the use of invariance.

The Byzantine Generals Problem

Host: I want to talk about the I guess the next paper which uh is the Byzantines general's problem. I think that's something that we hear about and we learn about when you're going through college and computer science and the name is great and I want to know the the story behind that problem.

Leslie Lamport: After I wrote that time clocks paper that was a tells you how to build a distributed system but assuming no failures and it was obvious that um you know distributed one reason for a distributed systems is you have multiple computers so if one fails you can you know keep going. in particular uh that was the problem that it was being solved at SRRI when I uh when I joined it but before I got to SRRI and I started working on that problem and I uh there's no notion of idea of you know what I should think about is you know what what can a failure do so I assume that you know the worst possible case that a failed process might do absolutely anything.

And I came up with an algorithm that basically would uh implement a state machine uh under that assumption and that the algorithm I came out with used digital signatures. Yeah. So that it used the fact that a faulty process might do anything but it could not forge the signature of another process.

Host: which just means that the message can be trusted that it came from a private...

Leslie Lamport: right so that you can relay messages and the people know can check that the relayed message is actually the one that was originally sent uh and so that a solution using that when I got to SRRI I realized that the people were were trying to solve the same problem. Uh but there are two differences. First of all, at the time I did this was you know 1975. very few people know knew about digital signatures and in fact I don't remember when the Diffy Helman paper was published but it was around 1975 and I happen to know about digital signatures because Whit Diffy who was one of the author two authors of that paper uh was a friend of mine and in fact at one point we were at a coffee house uh and he was describing these things that he said we have this problem of building digital signatures uh you know we haven't solved and I said oh that seems easy enough and uh and I sat down and literally on a napkin I wrote out a a you know the first digital signature algorithm.

It was not practical at the time because it it required basically something like uh you know 128 bits to sign one bit of the you know of of the thing you they're signing. It's not quite that bad because you know as you might think because you could use sign not a the entire dent document but a hash of that document which you assume you know people cannot forge uh

Host: the hash they can't reverse

Leslie Lamport: yeah you can't reverse you go take a hash and and you know you find some other hash that you know or some other document that satisfies that hash. But anyway, that's why I had, you know, digital signatures were part of my toolkit. Uh, so the people at SRRI didn't have that, but they also had a nicer abstraction of it. Instead of getting agreement on a sequence among the processes on a sequence of commands, uh they would agree have an algorithm for agreement on a single command and then that algorithm would be uh executed multiple times to and you know that was a nicer way of of describing uh you know what you're doing than than than my method.

So the first paper that was published uh use gave both the their original oh so but since they didn't have digital signatures they used a different algorithm uh and they had the property that to tolerate one faulty process uh you needed four processes whereas if you used digital signatures you only needed three processes. So the original paper contained both algorithms and so I was one of the authors. The other algorithm without digital signatures is is more complicated and the general one for end processes was really a work of genius. Uh it was almost incomprehensible. You just had to read in this complicated proof that uh you know for the arbitrary case of an arbitrary number of processes you need n pro for to tolerate n faults you needed four n processes whereas with digital signatures you need three n processes and the the algorithm for single fault wasn't hard but the one for multiple four parts was uh Marshall Peas was the one who did it and just brilliant uh Later in a in a later paper I uh I discovered uh a simpler proof one that was an inductive proofly proof that if it works for n minus one you know you it worked for n with 3 n if it works for 3 n * n minus one the original paper was uh you know the original one was just brilliant uh who would have discovered it anyway um so we published that paper and I realized that this was this the whole idea of Byzantine fault.

So the thing is that Byzantine well Byzantine fault is one that where process assume a process can do anything. Now I was assuming that you know processing can do anything because you know I didn't know what to assume but the people at SRRI had the contract for building a multiprocess multi- computer system for flying airplanes and so they were the ones who appreciated the need for solving processes that can do malicious things because they they really couldn't assume what it would do. And every time you would get an algorithm and you you'd see, oh, uh, well, this algorithm, you know, try to get an algorithm with three processes, you know, for one fault, you know, you'd find that, you know, oh, you know, this this works and it must be, you know, really couldn't happen in practice. And then you'd be able to find some sequence of plausible failures that would lead the algorithm to be defeated if there were a faulty process.

So you needed four uh and for some reason you know I thought that digital signatures was almost a metaphor in the algorithm that it should be possible you know since we weren't worried about malicious failures but but you know just things that happen randomly that there should be some way of of writing a digital signature algorithm that uh you know would have a sufficiently low probability of of failing but I never worked on that and nobody else ever did. So that those that algorithm was was pretty much ignored because digital signatures were very expensive in those days. I don't know what's being done now because you know computers are digital signatures are just computing and computing is you know is cheap. Uh but uh I remember at some point I happened to be communicating with someone who was an engineer at Boeing and I asked whether they knew about those results and he said yes when that he in fact uh was the one at Boeing who would read that paper and his reaction was oh we need four four computers.

Uh but at any rate I realized that this was an important result and it should be well known and I had learned one thing from Dystra. uh Dy, you know, one of the things I learned from Dystra, he wrote this paper called the the dining philosophers problem. And that paper got a lot of attention, but the dining philosophers problem, I won't go into what it is, but I think the basic problem uh was not particularly interesting, but it had a cute story to it. It involved a bunch of philosophers sitting around a table with uh some funny kind of spaghetti that it required two forks and there was one fork between you know each fork would be shared with two people and uh but and I think realized it was because of that cute story that that problem was was popular.

And so I decided that you know this this our work needed a cute story you know a nice story and I in invented Byzantine generals with the idea being that you have a group of you know for the for the one failure case you have four generals who have to agree whether or not to attack. uh and if they all attack uh they'll win the battle. But if only some of them attack or if even if three of them attack they'll win the battle. But if only two attack you know they would lose. But one of the the generals might be uh a traitor. And so how could you you know solve this problem? And so so it's phrased in terms of these generals having to communicate and decide whether to make the single decision whether to uh attack or or retreat. Um and you know I called it the Byzantine generals uh problem.

Host: I saw in your your uh notes about the problem that there was maybe a subset of the problem or a prior version that was called the Chinese general's problem or something like that.

Leslie Lamport: Oh yeah. that yeah I was uh there was a different problem that uh Jim Gray uh described uh as an impossibility result basically it's called the Chinese generals problem and I I won't bother going into what it is and so that gave me the idea of generals uh I actually initially thought of the idea of Albanian generals because at that time Albania was a black hole as far as the rest of the world was concerned. It was a communist regime, a part of the the Soviet uh block, but it was even more Soviet than the Soviet Republic and and and you know, more restrictive. So someone uh my boss said, "Well, you know, there are Albanians in the world, so shouldn't that so should have a different name?" And then I I realized that Byzantine there aren't any Byzantiums Byzantines around and that was the perfect name.

Host: So it's interesting to me in the story that because this isn't the first time the problem was specified but it was the first time that you had named it uh um gave it a good catchy name essentially and and uh you know added some additional results. What was it that you saw in that problem that made it interesting? Or rather like how do you know that a problem is worth putting extra time into?

Leslie Lamport: Oh, well this one it was because you know the it was obvious that people were going to be building that computers were going to fly our airplane fly airplanes and the reason in fact because was was that this was during the the time of the oil crisis in the 70s and that they knew people knew that they could build more energyefficient planes by reducing the size the the size of the control surfaces. But that made the plane aerodynamically unstable. Uh and a a pilot couldn't make the all the adjustments needed to, you know, to keep it flying, but a computer could. So it was clear the future was, you know, airplanes were going to be flying be flown by computers as they are, you know, today. uh and uh people didn't realize they thought that oh if you want to be able to tolerate one fault you just use three computers and they didn't realize that you know with arbitrary faults you need four and so that a really important result and that's why I believe that it it needed to to be well known.

Problem Solving & Paxos

Host: generally when you look at the problems that you are solving with your work Um, how'd you decide? Cuz if you're working at a company, you can decide based off of maybe the I guess the impact to the company like is it going to make more money or save cost or something like that. But I wonder in your work across your career um you know think about the bakery problem or some of your later work as well. How do you know it's it's so open-ended. How do you know which problems are uh the ones worthwhile?

Leslie Lamport: Throughout my career, I worked for private companies, you know, not, you know, not in academia or or for the government. Uh, and so some problems arose because of, you know, sometimes, you know, an engineer would have a problem and come come to me. And so, uh, you know, DIS Paxos, for example, was was a case of that that somebody actually wanted an algorithm to do what it did.

Host: You mentioned earlier Paxos and I know that's one of your your most famous uh works. Curious about the story behind maybe that paper and the problem you're solving.

Leslie Lamport: Well, the problem I was trying is exactly the same problem as I was solving in the the Byzantine general's work uh basically building a a fault tolerant state machine. But by that time it was you know the in the faults that interested industry were ones where failure meant that the computer just stopped the not not that it did arbitrary things. So uh the paxos is an algorithm for uh for for building fault tolerance systems for handling that class of faults.

uh and the people I was working at was which the it was the deck circ was in which I joined in 1985 and they built a uh one of the first operating systems that uh was a a distributed operating system. Uh so that um basically everybody had the they basically these are the people who had come from Xerox Park and had invented personal computing but they also had the notion of distributed personal computing and they invented the Ethernet uh you know for that. So they basically all of the uh computers in the building were on a single Ethernet network and shared a common storage uh and they had an algorithm for maintaining consistency of that storage and I didn't believe well they didn't have an algorithm they had an operating system with code that did that um and I didn't believe that what they what they did was possible. Uh, namely I I didn't think um well I forget exactly why I didn't think it was possible but at any rate I started you know trying to uh come up with a a an impossibility proof and start solidity proof well and an algorithm to solve this would have to do this and in order to do this it would have to do that and at some point I stopped and said oh this isn't a proof. It it can't. This is an algorithm that does it.

Host: You said that they had code but not an algorithm.

Leslie Lamport: Yeah.

Host: Um what do you mean by that?

Leslie Lamport: when most people sit down and start writing programs that you know they start by thinking in terms of code and one of the things I learned fairly early in my career I don't remember exactly when that back in in the days when I started writing algorithms people talked about people were calling them programs and I was probably calling them programs too I mean I remember then at some point I realized that that wasn't wasn't talking about programs. I was talking about al interested in algorithms.

Uh and an algorithm is something that's more abstract than a pro than than a program. U an algorithm can be you know a program is written in in some particular code. But an algorithm can be implemented if programs written in any any kinds of code. It's it's something that's that's at a higher level of of of abstraction. And of course I like that because abstraction is something I'm I was good at, you know, even without realizing that that's what I was doing.

Uh and so what I've spent a large part of my career basically from maybe about you know 2000 or so onward uh was getting people who build concurrent systems uh to not just write code but to have an algorithm. Now a system does lots of things but there should be some kernel of the of the program that's that's involved with synchronizing the different processes or the distributed system the different computers and that code you know is very hard to get you know the you know that correct so you you don't want to think in terms of code because static coding, you know, conflates, you know, a lot of issues that are irrelevant to the concurrency aspect. And so you should be thinking, you know, first get an algorithm that does that synchronization and then implement that algorithm.

Host: I was looking at the Paxos paper and uh some of your notes about it and I saw that um there's a there's an eight-year gap between when you came up with the algorithm and when the the paper was actually published called Part-Time Parliament is the name of the paper. Why why is there an eight-year gap?

Leslie Lamport: Oh, well the re the referees originally said well this paper is okay you know not terribly important but fortunately Butler Lamson realized the importance of the algorithm and together with the idea of you know I guess you can implement anything because it's implementing a state machine uh and you know went about proceed uh procilitizing uh building your systems you know using paxos uh you know and and thinking in terms of state machines and uh so you know I wasn't uh so the idea was getting out so you know I was in no hurry to publish so you know I just let the paper sit and eventually uh there was a new editor that came along and uh uh he said that you know I think the status of the paper was that it was just uh you know it had been accepted but had no and needed revision and uh so he decided that yeah let's you know to to publish it and uh it was eventually published with little some a few things that uh to take well to to mention work that had been done in the in the in the uh interim and what I got is uh got a Keith Marzulo uh you know to do that part for me.

Uh and uh so the story was that this manuscript this was that well the story about Paxos was that you know was a this happened you know centuries ago and you know this manuscript and uh I used that to the effect that you know when something you know the tales of something were I considered obvious and you know not interesting you know the the paper would say it's not clear how the Paxons what the Paxons did you know at this point but Um at any rate and uh so uh and Keith you know kept up the that idea that you know this was a you know a description of this ancient thing and and he wrote a you know a little prefix or a preface or something to to it and uh you know added maybe I think some uh references.

Host: I saw in your writing too when you were talking about presenting the paper initially, you even uh dressed up in like an Indiana Jones style archaeologist. Well, how did that go when you presented about this Paxos uh paper and algorithm?

Leslie Lamport: Well, I think the the lecture may have gone well, but uh I think nobody understood the algorithm where nobody understood the significance of the algorithm.

Host: It sounds like no one understood it except for Butler Lamson. What what did he see that made him unique? I guess.

Leslie Lamport: Well, he had a good understanding of building systems, you know, he really deserved his touring award. He was one of the original people at Xerox Park who were building distributed uh personal computing. He and Chuck Thacker, I think, were probably the two senior people, you know, in that lab.

Paxos vs. Raft

Host: I saw later there was a paper which describes a new algorithm which seems to solve the same problem. The raft paper. I was wondering if you read that and what your thoughts were on on that versus Paxos.

Leslie Lamport: The authors of that actually sent me a draft of the original paper and I looked at it and said uh I forget whether I said send it back to me when you have an algorithm or send that back to me when you have a proof. I I forget which one it was and uh you got the idea and they really they they did write you know add a proof in the paper or not. Uh yeah and I never read future later versions and someone whose judgment I value said you know had read it and said that it's basically it's it's the Paxos paper but no but with some of the the tales left unfinished by the Paxos paper uh by uh you know filled some of the tales filled in but they you know described it in a in a very different.

Hey, the basic idea of the what Paxos works is it's two phases and you're trying to implement a sequence of you know of decisions and it turns out you can do the first phase once for a whole it involves a leader. So um and the leader has to get elected. Uh so but it turns out that you can do the first uh phase once uh and you don't have to do it again as long as you have the same leader. Uh but it's only the second part that you have to do and then you have to elect the the new leader if a new leader fails and do the first part.

So think about it in those two phases. But the way people the way engineers you know like to think about it is well you do this you know you talking about the first part the the second phase you keep doing this uh until the leader fa fails and then you go back then you have to do this thing so it's explaining it in the in the opposite order uh and in fact you know when you start it from from fresh the uh you don't have to do the first uh uh the first phase you can you basically what what's done in the first phase could be just built in into the initial state but you know I think that that's the right you know of those two phases the way to understand it.

Uh but you know the raft people also had this idea that you know raft is better because it's simpler and I I must say that a lot of people say that uh Paxos is hard to understand and I don't understand why. I mean, I've explained it to some people in five minutes and they understood it. At any rate, the raft people said that one of the ideas were simpler because and they even have, you know, taught, you know, Paxos to one class and and uh the raft to another and they took and then yes, the people all the students said that yes, it was more understandable.

Uh the interesting thing about it though is that uh there was a bug discovered in raft and fixed but I believe the algorithm that they found more understandable was one with that bug. So uh made me realize that uh you know what most people you know what does understanding mean and for me understanding means you know you can write a proof of it but what understanding means for most people is warm fuzzy feeling and you know the raft description gave them you know more of a warm fuzzy feeling because you know you you know that that was seems to be the way you know programmers, you know, like to think about the the algorithm, you know, you know, the second phase, you know, first until, you know, you get a failure and uh but the way I describe it is one that helps you get a better understanding of why it actually works.

LaTeX

Host: So, yeah, we talked about a lot of your papers. I know one of your other uh contributions whether you knew it or not at the time was latte and uh building that and something that has impacted the entire academic community. What's the story behind wanting to build latte?

Leslie Lamport: Oh, that was uh very simple. Um, I was wanted I was in the process of starting to write a book and uh it was clear that tech was the basic uh type setting system that one had to use. But you know I felt that I would need macros uh to make tech do what I wanted it to do. And uh so I decide figured with uh been a little extra effort uh I could make the macros usable by other people. The system I had been using before tech it's called scribe and uh that really had the basic idea of scribe was that you describe the logical structure of of the document not and the and scribe will do the formatting. Well, scribe didn't do that great a job of formatting. Uh so, uh but obviously, you know, I like the idea abstraction that it's the ideas that matter, not the text that ma, you know, the the writing that matters, not the type setting.

And so, um, I actually at some point, uh, I met Peter Gordon, Addison Wesley, uh, I'm not sure what what you would call him, but he looks for, you know, books to publish. And, uh, he convinced me that I should write a book on it. And those days, it never occurred to me people would actually spend money for a book about software. But you know what the hell? And what he did was he introduced me to uh a typographic designer at uh Addison Wesley who was responsible for really for the typographic design that's in the standard uh latex styles. You know, basically I just did that in my quote spare time. You know, took me six or nine months or so. I I suppose the uh statute of limitations has run out, but I was really, you know, spent some time working on that when I was allegedly, you know, billing the time to some project that had nothing to do with it.

Writing, Thinking, and Proofs

Host: On the topic of writing, you have a quote that I really enjoy. It's if you're if you're thinking without writing, you only think you're thinking. And I was curious to hear your thoughts on what you mean by that.

Leslie Lamport: Well, it was really meant for, you know, people building computer systems. You have an idea and you think it's going to work. Uh or you have something that, you know, you think is something that somebody else will you want to use. Well, write a description of it. Uh there's an old maxim that I don't I heard uh that is you know write the instruction manual before you write the program. a great advice. Uh I did not do that uh with latte but it I definitely when I was writing the book and I discovered that something was hard to to describe hard to explain that needed to be changed and I made you know a number of uh of changes to it uh as a result of that but uh I didn't start at the beginning with the instruction manual.

Host: Why is writing conducive to good thinking?

Leslie Lamport: Because it's very easy to uh it's very easy to fool yourself. Uh I mean that underlies my uh my whole idea of of writing proofs. One thing I learned is that you had to write a a correctness proof of an concurrent algorithm. And when my algorithm was starting to get more complicated, the proofs started I started writ PhD in math. I knew how to write proofs and I was starting writing the proofs the way I would normally do. And I realized it just didn't work because there were just so many details involved and I just couldn't keep track of them and whether I had done it.

And so as a computer science know how to deal with concurrency uh it's hierarchical structure and so I devised this hierarchical structure where a proof is uh you know is a sequence of steps each of which has a proof and the proof is either a par well a proof is either a paragraph or a statement a sequence of steps each with its proof and that proof can be either a parag graph or a sequence of steps with its proof and you know so you break the whole problem up into these smaller pieces. So there's never any question of you know where is this coming from. You know you're stating that this step follows from you know this step this step this step this step and if it does not follow from that step your proof is wrong. The theorem might be correct but but means your proof is wrong.

Um well you know so I've discovered that worked great on writing my proofs of programs but I decided to really you know I also write proofs of theorems you know you know uh you know think proofs that are things that are you know more like ordinary math and I started trying that on them and I discovered it worked beautifully. So when I started to to try to convince mathematicians to write these proofs uh I started in one small seminar I went you know won't describe what it was about but uh and I I described this this proof through maybe uh 20 mathematicians or something their reaction shocked me they became angry I really thought that they might physically attack me.

So I believe that what's going on is that when pe I mean I believe that's totally irrational and when people act irrationally it tends to be out of fear and what I believe people are afraid of is that mathematicians are afraid of is that they're going to have to write their proofs to convince a a computer program and and in fact you know and I give it a one of those talks I gave you know I say very clearly this doesn't have to be you don't have to be any more formal than you do you can write the exact same thing proof you know but it's just a matter of organizing things and it's very simple you know hierarchical structure and then when you're using a fact mention that you're using that nothing about formalism or anything you know after I gave that talk someone got up and said I don't want to have to write my proof my my proofs for a computer program.

And in fact, it's more work doing that because the reason it's more work is that it reveals what you haven't said and that there's steps in there that you know, you may think they're obvious, but you haven't written them down. And if you believe something is correct but don't really if you if you think you know something but don't write it down, you only think you know it. And that's where errors come in. You know, that's where that one-third of your paper's errors can, you know, you know, uh, come in because it really makes you honest.

Career Reflections: Industry & Academia

Host: When I look across your career, I think you had a lot of contributions people might expect might come from academia, these papers and things, but you did uh all of your work in industry. Why did you not see yourself as a academic and more of uh working for industry?

Leslie Lamport: Well, I started out programming uh and I eventually got jobs where took me into what we now call computer science. At the time I never even realized that uh there was you know there could be a computer a science of computing. Uh it wasn't until you know maybe until mid to late '7s that I realized yes there was a computer science and know as a computer scientist. Um but it never seemed to me that like computer science was a an academic subject. At some point I, you know, had to make a choice between doing computer science and without calling it computer science or or teaching math at a at a university. And I I chose for random fairly random reasons to you do computer science. Uh so it you know for the first I don't know well till maybe the mid80s or something it just didn't seem to me that you know you know computer science was something that people needed to go to to a university to learn. And uh I suppose afterwards that I was sort of I guess I I just didn't think it would be fun teaching computer science. So

Host: I saw in your your writing you had a footnote that said somewhere that you you you felt like a a failure at some point because you you wanted to develop this grand theory of concurrency and you never discovered it. Um do do you still feel that way or what are your thoughts on that that footnote?

Leslie Lamport: lots of people who you know a large percentage of the people who were doing things like I was doing which is not a large number of people uh there's this notion that uh they're looking for the touring machine of concurrency you know the touring machine was this abstraction which really captured what computing was uh and they were looking for something that would be the you know the touring machine of of concurrent computing and you know nobody succeeded. I mean there are some people who think they've succeeded. Uh the patronets are are something that uh I guess I don't have time to to explain but uh there was a big it was big in the 70s. Uh, and I was actually surprised to think that there's still a large community of people doing uh, patriets. But what I now realize is that patriets and most of the things that people were doing was really language-based.

And I was never interested in languages. I'm interested in what the language is expressing. And you know I realized in some sense you know maybe I've realized what the touring machine of of of computing is state machines. Uh state machines are a little bit different the way I now describe them. They don't have commands. They just have a state and a and a next state relation. Uh even simpler than talking about commands and and and values and stuff. Uh and you know to me you know that's the uh that's the touring machine of of con of concurrency but it uh it it doesn't have the function that that touring machines offer because it it doesn't what touring machines do is uh describe what's you know what's possible uh and state machines can describe anything including things that are not possible.

Uh and and in fact uh the there's a good reason for that. Um for example uh when I describe a uh an algorithm I will talk about you know the values of a variable you know can be any integer. Now you can implement the program where you have any integer uh but that makes the but talking about you know computer integers would complicate things unnecessarily. The people have this funny idea that you know because something is infinite it's more complicated. They got it backwards. Infinity was introduced to simplify things. You know the first thing you learn is arithmetic. You're learning arithmetic with an infinite number of integers because if you restricted to a finite set of integers, arithmetic becomes much more complicated.

So you know the abstractions of mathematics uh which people find you know because they don't have the proper training in mathematics find you know difficult uh are really what's simplifying things and that's what you what you use this mathematics the state machine is described for me by me using mathematics that's the right you know the the most powerful way of doing But computer people and computer scientists and programmers are really hung up on languages and so they are looking for you know they invent all sorts of languages and they're all describable and in fact if you want to give them a semantics you would do it in terms of a state machine and they just think that this uh you know this language ES improves your thinking. uh it doesn't you may I mean there are reasons why you use computer languages and you don't write your your programs code in math and they involve basically efficiency but for understanding you know you can't build math you can't beat math and you know attempts to uh do it by something that looks like a programming language uh is is just the wrong way to to to deal when you're trying to deal with concurrency.

Host: When I look at uh everything that you've written and all the stories, there's these little anecdotes. There's things where you say things like you you never considered yourself smart, but you noticed that other kids had an awful time understanding things or yeah, there's a problem that you solved where someone else had difficulties, but you don't view your contribution as a brilliant one or anything like that. And that that uh doesn't connect with me because you've also won a touring award and done all these amazing things. So, how could that be that you, you know, just merely discover things and are are not smart yet you've achieved so much?

Leslie Lamport: Well, this general thing that, you know, psychologists talk about uh is that when someone is good at something, they don't realize how they're good they are at it because it's simple to them. There's the opposite one that uh people who are bad at something think they're better than they are because they're bad at it. Or to put uh a little bit more concisely, stupid people think they're smart because they're too stupid to realize they're not. Uh my the gift that I have is not in some sense raw intelligence. It's abstraction and it's only recently, you know, the last 10 or so years that I realized how much better I am at that than other people, most other people.

Host: At this point, you've experienced so much and you when you look back on your career. If you could go back to yourself when you just graduated college and give yourself some advice knowing what you know now, what would you say?

Leslie Lamport: One thing I've learned fairly early in my life is that I shouldn't waste time trying to answer questions that I don't have to answer. I don't think about, you know, what I should have done because uh that's a question that I don't have to answer.

Outro

Host: Thank you for listening to the podcast. It's a passion project of mine that I've really enjoyed building. Another passion project that I've been working on kind of in secret is building an ergonomic keyboard that I wish existed and I finally have a prototype. So, I'd love to show you what we've built. It's ultra low profile and ergonomic. and I couldn't find anything like it on the market. So, that's why we built it. I'll put a link to the keyboard in the description. You can take a look and learn more about the project there. We could definitely use your support. Also, if you have any feedback for me about the show, I'd love to hear it. Comments on YouTube have led to guests coming on like Ilia Gregoric and David Fowler. I wasn't aware of them until someone dropped a comment. Also, feedback in the comments helped me learn to reduce the number of cliffhers in the intros. So, your comments definitely make a difference. Please keep letting me know what you'd like to see more of in the show, and I'll see you in the next episode.


n5321 | 2026年2月27日 12:13