TDD, Where Did It All Go Wrong
n5321 | 2026年10月7日 19:47
Um, just a quick sort of show of hands, or a couple of questions from the audience: How many of you have actually seen a video of me doing this talk at any point? A few of you, okay. There are some nuances, and giving it live will be a bit different, but you're seeing the gist of the argument.
How many of you are using Test-Driven Development today? Test-first or test-after? Okay. How many of you are doing it test-first? All right. And how many of you are using mocking techniques and mocks? Okay, loads of you. Okay, very interesting.
Okay, this is me. Most of that is dull and uninteresting, but what I usually draw attention to is the last point on that slide: people like me come and talk at conferences to stand in front of you—we're not necessarily smarter or cleverer than any of you guys. And you know, I'd really encourage all of you to get up and speak at conferences if you can, contribute back to the community. That's all what I'm doing. I'm just a developer like you who decided to try and share some of my ideas with everybody else.
I kind of got forced into it because back when .NET first started, there were very few people who were the experts. And so, as such, you know, if you wanted to run a meetup or a user group, you had to kind of start up and do it yourself. But I'd encourage all of you to get involved if you can. It's a great way of personal growth for you as a developer in terms of learning to express your ideas more clearly.
That's who I work for; I work for Huddle. It's not really that exciting—we are hiring, by the way—but it's not that really exciting for you other than just to say we are a SaaS software business. I have nothing to sell you, right? You can't buy my services. So, all of the ideas here are based simply on my experience I'm trying to share with you, not on me trying to basically market anything to you.
What we're going to talk about today:
I want to talk about what I perceive as being the problem with how people are practicing TDD today and why there has been growing resistance to TDD.
I want to then look at rebooting our TDD practice, taking us back really to where we started from, and looking again at ideas like Red-Green-Refactor, clean code, refactoring, etc., to understand how those ideas should be applied, where we may have made mistakes in our practice that have led us to adverse ways of doing TDD.
And then I'll talk about some ideas that can help: Ports and Adapters, gears, look at ATDD, BDD, and mocks. We want to get through some of the later sections; we may go a little bit faster depending on time, so we'll try and focus on the early sections.
First thing I'll do is talk about the problem.
Now, I have been practicing TDD since about 2004 in London. The .NET community were quite early adopters of Test-Driven Development, just simply because we had close relationships with an XP group that used to meet in London and a lot of cross-fertilization of ideas. And so we were into practicing TDD for quite a while.
And I first gave this talk in about 2013, so I'd been roughly 10 years doing TDD. And I had come to the conclusion, after owning TDD-produced test suites for a while, that there were some problems with what we were doing. And you know, one of my things is that I tend to try, because as you become a software architect, to stay at companies long enough to see my own mistakes come back to meet me. It's a very humbling thing to do, but it's also how you gain kind of a lot of value from your experience. If you stay somewhere two years, quite often the problem is that you never see... you finish implementing something, you're very proud of this great work you've done, and you move off. And you never find out what were all the things that you actually did wrong that later developers whined about. So if you stay a bit longer, quite often you find what goes wrong. And I had written quite extensive test suites, and I found some of them were very difficult and very expensive to own. So that's going to talk about my understanding of TDD and how that shifted over the years.
Okay, so the first thing was back in 2005–2006, when I had first drunk the Kool-Aid on TDD. When I was practicing TDD, I found quite a lot of resistance to Test-Driven Development. A lot of people were saying, "We don't want to do this. This is a crazy idea." And I was very much, "You just don't get it. You just don't get what the rest of us get. You know, this is the best change in development practice I've seen in the last 10 years. You guys need to be doing TDD. You just don't get it."
But still, the people saying, "Well, I just... I'm not really sure about this TDD, this crazy thing." A lot of those guys were really quite smart developers who I respected. You know, these are guys who wrote really good code. Yet they were saying to me, "This TDD thing, it's not the right way to develop software." And you know, the more they said this, the more I just kept repeating the TDD mantras, as if like a child—if I just kept speaking slowly and carefully and repeating how TDD worked, eventually the light bulb would go off and they would appreciate the glory that was TDD.
But a few of them just seemed to resist. I remember one of them, a really smart developer who I really, really admired his work, said to me one time, "Okay, and I finally got it." And I was like, "Yes?" And he said, "The junior developers—they need TDD. It really helps them. I can see how it helps them write good code. But surely you don't really expect the experienced developers on the team to be doing TDD, right?" And I was really puzzled. I couldn't really understand, what is this guy getting at?
A couple of times, we were very proud of the fact that we produced some code that went into production with almost zero defects. There was no UI involved, making it a bit easier, but one time I worked on a backend service and that went to production, and for a year there were no reported defects. Worked exactly as expected. We were very proud, and we put it down that this had worked with TDD: we had a fantastic test suite, we knew all the scenarios we were going to operate under, and provably our code worked.
All I looked at this test suite, one of the things I noticed was that we had about three times as much test code as production code. And I can remember being a little bit, "Mmm, is this right?" So I went onto Twitter, source of all answers to all questions in life and universe, the source of truth and wisdom. And I said to Twitter, "Twitter, is this right?" "Yeah, no, no, we quite often have, you know, two or three times the amount of test code that we have basically as actual production code." I was thinking, "Okay, but it feels a bit puzzling."
One of the problems with this is you quite often got the situation where, you know, some project manager would turn around and say, "Guys, could we just skip the tests? We'd be done so much faster. We'd get out to market a lot faster. I know you guys like doing good engineering, etc., I really get that. But it's taking us an awfully long time to write all these tests. Couldn't we just ship the code? It wouldn't really matter if there were a few defects, you know, in production, right? Can we just not do the tests?"
And we'd say, "No, no, no, you don't understand. The way we write code is by writing tests. We can no more stop writing the tests than we can..." Let me collapse this all right into the standard mantra: "Tests are part of our practice for writing code. You're just telling us to stop writing code if you do that." And they'd go away and go, "Okay, fair enough, we get it, you know, you guys want to write tests."
But we were slower. I mean, in reality, we were, you know, 20% to 50% slower than we had been before by writing tests.
The other thing I noticed over time was with so many big test suites is that when we wanted to change something, we'd quite often break a lot of the tests. And quite often what we were doing was simply refactoring. We were going, "Oh, well, we've added some new features to the code, and now the shape of the implementation is a bit wrong given these new features. And I can think of a better way of organizing our code to make it more maintainable. So I'm going to refactor the way the code is organized. I don't want to change the behavior of the system; I just want to refactor the shape of the code so that essentially it's more efficient."
Wham! All the tests would break! And all the time, the tests that broke were the ones that basically were heavily mocked. Because these mocks were telling us exactly what ought to be happening in the code under test. They were saying, "Oh, first you should be calling this twice, then call this, and then call this and put these parameters on it." When we refactored something, we quite often changed how that worked. All these tests with mocks, they broke, and we had to rewrite all the mocks.
It was kind of puzzling, right? Because the promise was that we should be able to refactor our code, keep it basically fit and clean, we wouldn't get into a big ball of mud craziness, we wouldn't really see our code rot before us, because our tests would enable us to refactor, and refactoring kept our code healthy. But when we wanted to change our code, actually the tests were an obstacle to that change because they increased the amount of effort we had to put into making any changes. We had to fix all these tests that were breaking when we changed our implementation details. And that didn't sound right, because surely what Martin Fowler and Kent Beck were telling us about refactoring was that refactoring was changing the implementation details without breaking tests! Yet our tests were breaking all the time.
The other thing we began to notice was, I guess it's about 2009–2010, a whole new methodologies began to emerge that explicitly said: "Don't do Test-Driven Development, because Test-Driven Development slows you down."
So there's Programmer Anarchy. You guys know what Programmer Anarchy is? Uncle Fred George—Fred George essentially said (for those who understand *Spinal Tap* references), he said, "If XP is good, let's turn the dials on XP up to 11. Let's do Extreme Extreme Programming. Let's basically fire everybody who isn't a software developer: fire QA, fire the product guy, just have developers talking to customers. And what we'll do is we'll write really small pieces of code, we'll call them microservices, they'll only be about a hundred lines." So different schools of microservices—this is the original one: about a hundred lines of code. "And we won't have any tests because it's such a small amount of code you're bound to get it right. We'll put it into production; if you want to change, we'll just throw it away and rewrite it, okay?" That's what's called Programmer Anarchy.
Then there were other things like "spike and stabilize," even Lean Software Development, which became the real, you know, go-to approach of California and Silicon Valley startups, were saying: "Don't do tests, right? Work directly with customers, get it out there fast as you can, get feedback, because you may be building the wrong thing altogether and going to throw it away. TDD is a kind of practice you do later on when you're scaling. Up front, tests slow you down too much."
So there's this whole pressure saying TDD is a bad idea because TDD really slows you down and you're better off just getting feedback.
You guys especially over here: Duct-Tape Programmer. So the duct-tape programmer was this really annoying guy, right? And every organization has one. And he's the guy that business or the product team actually love, because he just gives them solutions as fast as he can. And as a requirement comes out, they say, "Let's give it to Bob. Bob will just get it out the door really fast." And the rest of the team are going, "But this is crazy!" The engineering team going, "We hate Bob! Bob's engineering is rubbish! We spend all this time maintaining the stuff that Bob's written and having to rewrite it!" And it's really annoying because Bob gets all the credit, right? From the product guy, the guys going, "Bob's fantastic! He's so receptive to our ideas, he pushes stuff out!" The rest of you are going, "But it's so hard to maintain, Bob's code is terrible, right?"
But they love Bob, because Bob delivers for them. You're still writing your test suite, and Bob has shipped the code and it's running, okay? We had a few of those, and it used to be really annoying. It used to be really annoying that we were still doing our best engineering practices, and these guys were shipping code—and it was awful code—but they were shipping, right? And so they won the credibility, not us.
Sometimes we'd come back to our tests when they were breaking, quite often because we were trying to change implementation details, we'd realize that we had no idea what the test was doing. Anyone have this experience? You go back and find the test is red, and you're looking at it and you're going, "I have no idea what this test is for, and I have no idea whether the fact that it's red is good or bad at this point in time, okay?" This test seems to add two numbers and come to a result which I can't possibly understand how they could have originally ever got that result, and why this test was originally passing, and now I've refactored it and it's red—surely just because the test is wrong?
Lots of tests just don't really give you any information, especially the ones that are like swarming with mocks, and you're kind of like, "What is actually under test here?" Anyone have the experience: I'd look at the test and then finally realize nothing is actually being tested apart from the mocks, right? Someone's written this thing around it so heavily, you go, "There's no real code actually under test here, okay?" And that experience basically began to drive us mad. We're like, "Well, I no longer know what this test is for. The best thing for me to do is just to delete this test because it's broken, I don't know why."
Made to my shame, a team I worked for embraced FitNesse, and we sat down and produced basically these living tables in HTML, and they drove... used them to drive our software. And eventually, you know, these little tests would show up, the page would show up green, so that we could say, "Yes, these high-level tests that are written in natural English so the customer can understand them, these tests basically demonstrate that our software does exactly what we say it does on the tin." And SpecFlow, Cucumber—these tools are all the same kind of thing: the idea essentially that I will write tests that the customer can just flip through the HTML pages, see the requirement we've implemented, feel confident that the system does exactly what he specified, because he can press the button and see it all go green. "Ah, the software runs exactly as I expected."
The developers basically hated writing these tests because they'd start out practically... they'd be red in the beginning of the iteration, so they'd be red for a lot of the iteration until eventually they would go green when you'd done the implementation. And that meant essentially your FitNesse test suite was quite often—well, correctly—red much of the iteration, so you could never really be sure if the tests were red, is that because it's yet to be implemented or is that because essentially we've broken it? And so the thing that happened was a lot of time the tests were red because we just assumed, "Oh, they're red because not implemented yet," when in fact they were broken. So they spent a lot of their time as broken test suites.
The customers never really looked at this stuff, right? We'd built these customer-facing tests, but it was very rare to get the customers to actually come along and flip through the HTML pages and say, "I'm so impressed, guys! I can see exactly the requirements that I want, I can see them written in a natural language, and I can see that it works, right?" Occasionally we'd try and sit them down, make them look at the screen, and they'd go, "Maybe send me a calendar invite, guys, and we'll talk about this at some point in the future, right?" And they were completely uninterested in seeing these test suites that we'd written to prove to them that the code works correctly.
And the check-in gates became more complicated, because as well as passing developer tests, you had to run all these FitNesse test suites before we allowed you to check your tests into the code base. They took a while to run; they were slow. And that meant essentially you'd go for making a cup of coffee and come back, essentially because, really, you know, that's you... then you'd lose track of where you were, what you were doing. And the whole flow of your coding session was broken, because now you're in that classic developer situation where it takes you 20 minutes to really get your head around the code that you're working on, and then you're in the groove and you're working on it, but now a forced interruption to our process is going to mean that you stop working effectively during that time period, okay?
So it became a really inefficient way of doing it, and the developers just really hated doing it and they started to ignore them, right? "If we just basically keep going and checking code in, eventually someone will notice the tests are all red and go break in, sort of thing, and then we have to basically down tools and fix them. But until then we can keep moving, right?" And then what happened was there would always be plausible deniability, right? "My test didn't make that go red. That change wasn't my change, that was Dave's change."
Then we used to have somebody appointed on the team whose job was to essentially determine who had broken the tests and made them go red, and then that person would essentially apportion blame for the person that had to go and fix it, right? The fact that we actually had a step in our process which essentially was "somebody decides who's to blame for the tests being red today and then they have to go and fix them" implies that no one really cared about those tests, because they'd reached that point of being red the whole time.
Really, the developers couldn't see any value. They constantly resisted writing these tests. And one thing I've learned over time is that, you know, if an idea actually really works, after a while your team will be quite happy behind it and you'll get some supporters of it. If your team is constantly pushing back saying, "We don't see what we're getting from this, I don't see the value in the tests that you're forcing me to write," then maybe you need to think again about what you're doing, okay?
And about a year after I gave this talk for the first time, DHH, the creator of Ruby on Rails, posted this blog post: *"TDD is dead. Long live testing."* And a lot of what his complaint was: "I am being forced by people to say that Rails is wrong, because Rails doesn't work the way they think test-friendly frameworks should work." So he said, "You know, I have the Active Record pattern in Rails. (For those who don't know Active Record, effectively the data persistence methods are actually on the object that you're persisting, so your user object has persistence methods on it.) I see people tell me this is no good because they can't cleanly separate their domain object models from basically their testing framework." And he said, "I just used to test... like write my unit tests and I talked to the database, and I'm now told that's wrong as well. And I said I used to like TDD when it first came out, but now I'm constantly being told I'm doing TDD wrong. Then I'm just giving up on TDD, you know? TDD as far as I'm concerned doesn't work. I'm still going to do some testing, but I'm quitting on TDD."
And of course, such a stir that Kent Beck kind of responded saying, you know, "RIP TDD." It was kind of a sarcastic post, which may not have been the best choice when Americans are amongst your audience, but he kind of said, you know, "Oh, well, here's all the things I'm not going to do now I don't have TDD anymore." So he's kind of defending TDD. And Martin Fowler actually produced this series of dialogues, mediated them between Kent and DHH, talking about TDD. DHH went and saw an original version of this talk I gave at NDC and was kind of saying, "Yeah, I think you and I are on the same lines." And they came to a pretty similar conclusion as the rest of this talk. So you might go away, if you like the ideas here, and go and see Kent and DHH having this dialogue mediated by Martin about what everyone thinks happened.
But the question is kind of, where did it all go wrong, right? What went wrong for TDD, and can we fix it? Can we change our ways?
Anyone know... anyone ever hear of an English footballer called George Best? No, mostly... very, very famous in the 1970s. And the story goes that a bellboy at the hotel enters George's room in the hotel. George is lying on the bed naked with a couple of supermodels. The bellboy is wheeling in the caviar and champagne that he is going to give to George and these naked supermodels on the bed, and the bellboy looks down at George and says, "George, where did it all go wrong?" And so treat that sentence in that kind of light, right? I am a fan of TDD. What I want to talk about is, you know, where did it all go wrong?
Okay, so what I want to do really is go back to the beginning. Let's go back to the beginning of TDD, and I want to explain to you: the premise of my argument is TDD has answers to all the problems that we have been meeting with TDD. We have just ignored the fact that basically the original approach to TDD as described by Kent Beck works. Later patterns and practices we have loaded onto TDD, in many cases, were a mistake and have led us away from what was originally intended.
So I would recommend: who's read this book? (*Test-Driven Development: By Example*). Okay, many more of you practice TDD than have read this book. Who read some other book to learn about TDD? Who was told how to do TDD by somebody else that worked near them? Yeah, probably more of you, okay.
I would recommend reading this book if you practice Test-Driven Development and you want to understand how it works. I recommend reading this book. When I began to struggle with TDD, I went back to read this book again. I'd read it originally, but I went back to read this book again. And one of the things I learned was that Kent had an enormous amount of wisdom about TDD, having practiced TDD for a large number of years before he wrote the book. And many of the problems I was encountering with TDD, Kent had already encountered, understood, and were discussed in this book. And this book, although it seems old, covers all sorts of things like mocking in TDD, which many people think are much later developments—they happened a lot later in the lifecycle; they're all in here, right? This book is probably the only wisdom you need, but I don't think I understood much of its wisdom until I'd been doing TDD for a certain amount of time, and I was then able to understand what Kent was trying to tell me, okay?
I think the most important thing to take away from an understanding of that book is this: **When you are practicing Test-Driven Development, do not test implementation details; test behaviors.**
I'll explore that sentence more as we go through this talk. That is the key thing that people get wrong today.
So what happens is, in a classic modern TDD cycle, someone says, "I'm about to write a method on a class. I will write a test before I write that method, and that test will govern whether that method succeeds or fails." And so the trigger to writing a test in TDD practice is essentially adding a method to a class.
That is the wrong thing! That is not the ethos of TDD.
The trigger in TDD for creating a new test is that you have a requirement you want to implement. It is nothing more; it is nothing less, okay? There is something I want the software to do; I will write a test to do that, okay? I want basically the software to add an amount to this customer's bank account, essentially, right? I want the software to upload a new file to the user's collection of files, right? That's what drives a test in TDD. That's what you are writing tests for. You're not writing tests to say, "The `Foo` method should take the two numbers, multiply them, and return the result," okay? That is an implementation detail. There is no requirement you received—unless you're writing a calculator—no requirement you received that gave you that. The requirement you received was something else much higher up, right? And that is where you write the tests.
What you want to be doing is focusing on testing the API of your software, okay? When we talk about the API, we'll talk carefully about that in a second, but we mean: What is the contract your software has with the world? What does it offer to do for people who are basically consuming your software? That is the stable idea of your software. It will not change that rapidly. You may add new stuff to it, etc., but it's a stable requirement you have. How you implement that requirement is unstable: you may get better ideas over time, and you may wish to change it. But what your software offers to consumers is the stable contract, and that is what you should test.
We don't necessarily, by the way, mean by this an HTTP JSON API, right? So I'm conscious now that when I use the word "API," everyone imagines, you know, REST, HTTP, JSON, GraphQL, whatever. What we mean simply by an API really is just the publicly exposed surface area of your module generally that somebody else is consuming. So in, you know, the classical description of a module: a module has basically a set of exports—public-facing things that other consumers can call, that other software can call—and internal implementation details that you've hidden away. These are classic notions of information hiding, encapsulation, right? Your module has a facade, something on the outside that essentially represents and hides all those details away. And it is that external sort of exports that you are interested in testing, not the implementation details which tell you how you are implementing that, okay?
What happens is we should get some use cases or a story in an agile environment, and that says, "Here is the requirement I want you to implement. Here is the story, here is the use case, here is the thing that your software is supposed to be doing." And you write tests to say, "Right, I'm going to write tests to prove I can do that." And generally speaking, there are good models for this like Given-When-Then: Given that my account has £100 in it, when I add another £100 to it, then I have a total of £200, right? That is a kind of classical requirement you're trying to implement. Now I may end up using objects to represent accounts and that kind of thing, but what I'm doing is saying my requirement is the behavior of my account, perhaps when I add or remove things from it.
**The system under test is not a class.** Too much TDD practice focuses on essentially writing tests for classes, okay? The system under test is really the exports for a module, its facade. When people talk about this phrase "unit testing," at the time that people were basically coming up with this kind of idea, to them in classical testing paradigms, a unit test was a test of a module, right? It's kind of a black box where effectively you talk to basically what's exposed.
Now, there's a certain line between classes and modules where effectively people can say, "Well, a class is a module—it has information hiding, encapsulation, etc." I appreciate that. But the definition of "unit" has become too much fixated on being a class. People talk about isolating that class and isolating the class's dependencies, right? Now, it leads into all sorts of problems, because then you start mocking heavily, which leads you into what I call over-specified tests where essentially you know all about the internal implementation details of that class through your mocks and how they're called, and essentially you've hidden no implementation details!
So you want to focus at a much higher level. You want to say, "I have some public things that this module does; that is what I'm going to test. I have some implementation details, and I'm not going to test those, okay?" And we will explain how to get there. TDD explains exactly how to get there. This may sound complicated, but there is an easy route to get there, okay?
So obviously, you may actually be driving a class, because if you're writing say C# or Java, you can only drive classes. But that class may tend to be a facade or a command, etc., some kind of interface to the actual module underneath. You don't want to write tests to cover implementation details because if those change... you want to write tests for your stable contract, which is essentially your exports from your module, your API.
Refactoring is the key step which enables you to achieve this goal of separating between things that are implementation details that you don't need to test and things that you do need to test. There is a get-out clause, and the get-out clause says sometimes it's helpful to probe implementation details that are difficult to understand—you can do that, but with test-afterwards; again, we'll come back to that.
So a lot of these ideas... I want to give a bit of proxy to Dan North. Dan wrote a post about what he called Behavior-Driven Development. Now that's turned into something else really, though it has the same origin point, the same root. This is the root of that. When people talk about Behavior-Driven Development today, they mean something else entirely; we'll talk about that later. But in his original post, Dan said, "I've often found when I go to basically implement TDD in places, people are very confused about how to do TDD because of the word 'test'. The word 'test' throws them. And if I change that word to 'behavior-driven'—developing your system using behaviors—people do it correctly." This insight was very old, you know, around 2005, etc. And he's saying exactly what really I'm going to tell you in this talk: people misunderstood what Kent was talking about, probably because Kent used classical testing terminology like "test", "unit test". He mentions it just a couple of times in the book; generally, the phrase Kent Beck uses is "developer test", right? Not unit testing—developer tests.
But Dan basically effectively recognized the same problem I'm recognizing here, which is people were driving tests by saying, "I need to test this method." They should have been saying, "I have a behavior in the system, and I want to use code to basically prove that that behavior works in the system, okay?"
So in case you think that's just a Dan thing, Kent wasn't talking about this... Kent Beck originally, Kent talked about behaviors: "Behaviors would demonstrate that code would compute a report correctly." So he's saying the object of TDD is to test behaviors in the system, okay?
Here is a classic thing that you are testing: "Add amounts in two different currencies, convert the results with a given set of exchange rates." We are not testing a method on a class; we are testing a behavior of the system that we can add two different amounts in two different currencies and then see basically the result in one currency. "We are multiplying a price per share by number of shares to receive an amount." Okay? The behavior is the requirement. These are TDD requirements as described by Kent Beck originally. That is what your test should be looking at, not "I need a method on my class for exchange rates." That's nothing... that's not... nothing to do with TDD.
So the idea behind Test-Driven Development has always been that we are essentially saying: What should the API of our code look like? If I am an external consumer of this code and I come to call it, what kind of API would I find useful? You're testing as much the usability of the API—how you interact with it—as you are essentially designing that public-facing API as anything else, okay? So you tell yourself a story about how this API should look if you were a consumer. The test becomes the first consumer of your code that essentially says what should this look like and how should it work, right? That is the objective of Test-Driven Development.
So Kent uses "unit test" in a very specific way, right? And simply for him, when he uses the phrase, all he means is that you should be able to run all your tests together as one suite without one test being run impacting the other test.
**The unit of isolation is the test.**
And a lot of mistakes have been made because people believe the unit of isolation is the class under test. It is not, okay?
So all these people mocking because they say, "Well, you've got to have the class isolated when you test it. These all have to be mocks. You can't interact with anything else. You've got to mock all the class's dependencies, right? Because otherwise it's not a proper unit test." You are wrong! The unit of isolation is the test, not the thing under test.
One of the problems is, you know, people were using definitions from classical testing. And the thing is essentially unit tests referred to this idea of a module, right? This is why it's a confusing phrase. Because when people talked about unit test, they were saying: take a module—so say, in .NET terms, an assembly or something, right?—take that and then essentially test it on its own. Then an integration test is I've got a couple of modules talking to each other, and then essentially a system test is talking... all modules talking to each other. So when Kent Beck used this phrase "unit test", what he was really talking about was a module being tested, okay? It's an unfortunate phrase, because people then said, "Well, the unit is a class," and came up with a whole set of practices which basically have been negative towards our TDD practice.
You may choose to avoid using certain interactions in testing, right? Databases, file systems, those kinds of things, right? But the reasons Kent gives for that are slightly different from the ones you think. It's not about isolation—it's about isolation only in the sense that if I touch a database, one test may impact another test because if I add some rows in the first test, when I read them in the second test, that may be impacted basically by the order they run in. So I have to make sure that between these two tests, I'm isolated from that kind of problem. Which may mean I choose to use something like an in-memory database and we tear that down rapidly and provide a fresh one for each test, right? But it's the isolation of the test that is driving that, not the isolation of the code under test.
Same with file systems, everything else. And the other issue is speed, right? Tests have to run fast; otherwise developers don't get fast feedback. So really, I want my test suite to run in a couple of minutes, and it's quite possible if I'm talking to real hard dependencies that becomes hard. So I may wish to essentially swap those out for the purpose of speed. But those are the reasons for it; there are no others. If there is no shared fixture problem, it's perfectly fine in a unit test to talk to a database or a file system. And that happens all the time in Kent Beck's testing. So people telling you, "You can't talk to a database in a unit test," they are wrong, okay?
Yeah, okay. Focusing on the methods creates tests that are hard to maintain, okay? Because essentially, when I look at that test, what I see is somebody describing how you interact with a method. That's why we get this problem when we come back to our tests: they're very hard for us to understand because we're kind of going, "Well, what does this test doing? Well, we've got a method in isolation, so it's very hard to understand what it's doing." If the test tests a behavior, it's very easy to understand what that test is doing, because it says, "This behavior is asserted to be true in the system." Now that test is red because I've changed another part of my code, I can either say, "Well, that should still be true," or perhaps, "No, that isn't true anymore. That behavior's been modified now by the new requirement, which means this behavior has also changed and I've just seen an impact of that behavioral change."
And when we have implementation details exposed—in other words, essentially all those things that make up how our API actually works—that's when you get this problem effectively of being hard to change, particularly if you use mocks. So when I use mocks effectively, when I say that's just a class under test and I know about all the calls out to mocks, the problem is when I change something that changes those, all the basically assertions break. And if I essentially therefore know too much about the implementation details, then I've coupled my tests to my implementation, and I can't change one without changing the other.
All right, so let's talk about how we get there, because the way we do this is really simple, and it's Red-Green-Refactor.
Who thinks they understand Red-Green-Refactor? Okay, great.
Who does Red? How many of you make sure that essentially... okay, very few of you.
Who does Green? Probably most of you.
Who then always does the refactoring step? Some, okay.
Refactoring is the key thing, right?
The test that doesn't work: okay, this is important. There are no tests for tests. I have to demonstrate that my test will fail in the absence of correct implementation, right? I've stumbled across this a few times, had the problem where essentially I had a test that would always pass. And so, you know, essentially until I got that red phase, I wouldn't prove anything by writing my code and having it pass again, because it was always going to pass anyway.
Green: Make the test work quickly, committing any sins basically necessary in the process. This is really important. What I'm doing is making the test pass by the quickest process possible, okay? A good solution here: cut and paste code from Stack Overflow, right? This is where the duct-tape programmer wins, because you're afraid to do that and he's not afraid to do that, okay? Now, make good code... but it's three steps: write a test, make it compile and see that it fails, make it run.
And then the third step is refactoring, okay?
When you write your implementation of your test and you go to your green phase, what you are seeking to do is understand how to solve the problem. You want to basically, as fast as possible, get something that satisfies the behavior: "I want to add two currencies in different amounts, and I want to basically look at the end result, right, in one currency." What I want to do as quickly as possible is figure out how to do that in my code. Just write a good transaction script: one line, with one line after another. Don't try and basically put it into lovely classes; don't try and essentially put patterns in there. Just write line after line of code until it works, okay? And that code can be dodgy if you want it to be—probably should be!
Your goal is to, as fast as possible, get that test to go green because you have managed to understand how you will implement the requirements. At this point, you are trying to determine how you will implement the requirements, and you can figure out the answer to that by going green. Go to Stack Overflow, copy some code, right? You are explicitly enjoined in this step to write non-well-engineered code. You want to be the duct-tape programmer! You want to be that really annoying guy on your team who writes bad code. Emulate him! Do what he is doing to get it out the door fast. You will now be keeping up with him. You've added the test, but he has probably done it manually a few times; you are moving at the same rate, okay?
As Kent Beck says: **For this brief moment, speed trumps design.**
I want to get there as fast as I can, okay? Kent's point is essentially you can't do two things at once easily: you can't both understand the solution to the problem and engineer the code, right? One of two things will go wrong: you will either over-engineer your solution beyond what the test actually requires, or you'll get some kind of analysis paralysis, right? But the first thing you do is solve the problem with the code.
Okay, so when we're getting our clean code... okay, so we believe in clean code; we don't want to have duct-tape programmer code all day long. And we're going to do that in the next phase: the refactoring step. And to produce clean code... so for Kent, basically the key driver is duplication, right? You see duplication, you remove that; that starts to re-engineer your code well. And that works there.
You can add into that, I think, Fowler—Fowler's book *Refactoring*, I'll talk about that a bit later: look for code smells, okay? Look for smells in your code that indicate there are better ways of doing this.
And Joshua Kerievsky's point is: if you want to think about using design patterns, this refactoring step is when you identify them. Because now you have a clear idea of what your solution is, you can tell whether a pattern is appropriate or essentially whether you're using patterns inappropriately as a solution, okay?
When you refactor—and this is the really important part—**you do not write any new tests.** Refactoring is a process of safe moves that let you essentially change the design of the code; they do not change its behavior. Your behavior is covered by the original test. The original test tells me I'm adding two amounts in two different currencies and getting the final result. The code I'm implementing produces that result. If I use safe refactoring steps to get there, it doesn't matter how much I change implementation details; I am still covered by my original test.
Code coverage tools help you here, because you don't want to introduce, for example, new conditionals, speculative code at this point. When you do that, you need to have new tests to cover those new behaviors, because generally you've got a conditional because you have a new behavior that you're trying to produce. But do not write new tests at this point!
This results in two things:
One, you will write less tests, so you will move faster. The refactoring step shouldn't take you so long that you fall miles behind this duct-tape programmer, right? You will keep up in the way that the lean software guys and your managers want you to, because you are not saying, "Oh, I'm creating a new class here, I'm doing extract class because I believe that essentially these things should be a class in my implementation." Don't write a test for that class, right?
If you have a language that supports visibility—for example, Java or C#, and you've got things like protected, internal, private, public, right?—in C#, for example, most need to be internal and never exposed. And as such, the test should not be coupled to them. You should be at liberty to change their implementation whenever you like, and no test will fail because no test is coupled to their implementation.
And this is key to understanding this, right? In the refactoring step, do not write new tests! You're now introducing a couple of classes. If you basically keep thinking, "Mmm, I need to introduce a public class," the only smell you may be looking at is that essentially in your solution there is some new thing that is also part of your API, okay? That will itself be exposed as part of the API of your module. In that case, that's probably an appropriate time to think about a mock for use in your test, and later go and implement that other part of your API that you're currently missing, right? But focus initially on trying to avoid writing these extra tests.
It says, right: dependencies are the key problem in software, right? Coupling is a worse enemy for you than DRY. Coupling is what kills all software. So struggle to decouple your tests from your implementation details. Focus on your public contract, your stable API, and leave your implementation details free of tests, and avoid heavy mocking.
This allows you to meet the promise of refactoring: you will refactor your code and your tests won't break, right? I've seen this work. Since I changed this practice, it's been very easy for us to refactor implementation details, to figure out new ways of doing things. We changed the whole way we did, for example, modeling an aggregate without any tests breaking, because the behavior hadn't changed; our choice about how to implement it had changed. So test behaviors, don't test implementations. (These mics are annoying if you wear glasses, right? Okay.)
Keep internals internal. Don't make these public to test. If you find yourself making things public in order to get them under test by implementation details, ask yourself if that genuinely is a part of the contract of your module. If it's not, don't! It's just implementation detail; why did you surface it for testing? If you are in .NET land and you use the assembly attribute `[InternalsVisibleTo]` and list your test assembly so that you can get your internals tested, do not do that, right? Your internals should not be visible to the test because they're not under test, okay? So preserve your information hiding; have a public API and internals.
Right, this is Fowler's book *Refactoring*. Actually came out before Kent's book on TDD, and it talks really about code smells and says here are some safe steps to be able to effectively refactor basically your code when you see these smells. And one of the key things to understand is, you know, very explicit, okay: we are changing implementation details, not public interface.
Anyone ever use the DevExpress tools as opposed to ReSharper's tools for refactoring? So DevExpress had a set of tools which basically had a code plug-in for Visual Studio, and they used to offer refactoring. And DevExpress's tools... some of the guys that worked on them were very clear, right: you cannot refactor a public method on a public class. They wouldn't let you do it! Because that's a change—that's not refactoring, because you don't know the impact of that. Now later they said, "Okay, you can press the button marked 'Unsafe Refactoring'," because everybody else lets you do it and we're being killed in the marketplace. But also because they said, "Well, you may say to me the only consumers of this public API are within the span of the solution, and therefore essentially I understand who all the callers are and I can then refactor them." So if I'm doing a rename, I know who I'm going to break. But the goal was originally: don't do this, right? Don't change... refactoring is not changing your public interface; it's changing how you chose to implement that. It is a safe, small change. There are recipes to do them, that's why tools automate them, and they're safe in the sense that your code is not basically going to... it's not going to negatively be impacted if you extract a method; it's a safe change, okay?
Here's the list of code smells. It's not really what this talk is about, but those are the kind of things that you're looking for, and those are the things that drive you in your refactoring step to do things like extract a class out of this transaction script body, extract a method out of this body, right? Look at basically where you've got Feature Envy, whether you're doing some kind of dot syntax, etc., right?
And Joshua Kerievsky in his book (*Refactoring to Patterns*) said there's a problem with patterns, and the problem with patterns is people are getting pattern-happy. They see patterns everywhere and start applying them, and in many cases what they've done is applied them in the wrong context or made their software over-complicated when it didn't need them. What Joshua Kerievsky said is: the ideal step for applying patterns is the refactoring step. Because at this point, I know what the solution is—I've written one! It's an awful solution, because I was told to make an awful solution! And at this point now, I can see I can clean it up by making a pattern. I can see that actually this is... which is the Command pattern, I need that here; this is the Template Method pattern, I need that here. But I know that now because I've already implemented it once badly. And even in Kent Beck's original book, he identifies patterns you may want to refactor to, so this is part of original TDD, it's not new stuff, okay?
So what about Ports and Adapters?
This is unfortunately the common way that testing occurs in a lot of organizations:
Before we release the software, there's a big, big wave of manual testing going on. Everyone's getting into the system, running manual scripts, etc., doing some kind of conference releasing, right?
Then below that, there's some sort of scientist who has big suites of Selenium tests that they run against their software. Not so bad? Funny of you, okay.
And then we have integration tests, and at the bottom we have a few developer tests, right?
This is a bit of a disaster! Why?
Well, first of all, this manual testing is really expensive, unrepeatable, so you really want to automate your way out of that situation.
This UI testing one is very problematic, and I'll tell you why: if I change my UI to give it a nice classy new look and feel that's much more, you know, 2017 rather than 2008, right, I may not be changing any other behavior of my software whatsoever. They do exactly the same thing that they did yesterday; I've just cleaned up the UI. But these tools all break, okay? But I'm driving them to test the behavior of my system, but the tools all break because I changed the way that my widgets work, okay? So they're very fragile, brittle, and slow. They also run really slow—quite often you run them overnight, and then you get the blame game in the morning: someone's job in the organization is to determine who is to blame for last night's tests failing, right? If you find yourself in the situation where you have somebody whose job every morning is to determine who's to blame for tests failing, you are doing it wrong! I know that because I've done that, and that's wrong, okay? All the mistakes I'm telling you about, right, I've made those! Don't worry, I'm not going at you because I think I'm smarter than you; I only know this because I've made every single one of these mistakes, right?
So what you want is what's called the Testing Pyramid: the majority of your code should be tested by your developer tests; a small amount of code up here should test your widgets work correctly; and in between, there's a certain amount of integration testing where the two get hooked together. There's a bit of a scam in the middle; I'll come back to that later, okay.
So this is a Hexagonal Architecture. It's gained popularity recently. People like Bob Martin called this the Clean Architecture; some people call it an Onion Architecture. There are various names for this architecture, various attempts at co-discovery of the same thing, all right?
What happens?
In the middle, we essentially have Plain Old C# Objects, Plain Old Java Objects, Plain Old Python Objects—anyone here ever heard that phrase "POCO"? That's also slang for the police in the UK, so be careful about using that word, but yeah! But these things are devoid of any technology concerns, so they don't know how to talk to the database, they don't know how to talk to the web; they're just basically completely plain old objects without any technology concerns.
On the outside, we've what's called the adapter layer. The adapter layer knows how to talk to the outside world; it has all the technology concerns. So our primary ones over here: things like basically an HTTP adapter, so something like basically your web framework—Web API, Django, Sinatra, effectively, or Flask, MVC, etc.—or a GUI adapter, so like Web Forms, or, you know, JavaScript, like that, right? And they talk basically to your domain. Over here, we've basically got a database that your domain uses to persist itself; over here, we've got essentially the idea of things like sending messages, things like RabbitMQ, or essentially that kind of thing, right? Or emails.
In between is what's called the ports layer. So the ports layer literally is like a port in an operating system: it is a door. It's the door which essentially allows one boundary to be crossed. Here, it's the door that essentially says, "Well, my HTTP code needs to talk to my domain; it would go through this door," some kind of facade layer which says we talk to it. Over here, it's saying, "My domain code needs to basically persist to a database, so I have a door that lets me talk to databases," some kind of repository model perhaps.
Now, okay, I've got a few five minutes more, and I'll just run through these quickly, right? Apologies for running over at the start, okay.
So the port here is the ideal point to do your... for your tests! The tests talk to that port layer, not to the internals here. Because the port is how everything communicates with your code, and also how everything your code raises—like events, or calls the database, okay? So the port is the point for you to make your calls.
Integration tests are really just checking that essentially you can talk to the database correctly, you're configured correctly on this adapter. Don't test adapters themselves; they're usually third-party code. And a system test just checks that you can call from the outside in, right? And that is essentially how you do it, okay?
Make two quick points:
I'll talk very briefly about gears. So you may say that's all great, and testing the API is fine, but I sometimes like TDD to help me understand how something's working, and I want to basically drill down using my TDD into the details, okay?
So Kent Beck has this notion of gears. How it works: it's notionally five gears.
In fourth gear, it's the practice you just described: you're testing the public API, and you're getting basically the full script out, and you're refactoring that, right?
In fifth gear, it's kind of like the implementation is very obvious, so actually that refactoring step—when I write my code, rather than write the sinful code, I just write the correct answer, so it's a regular refactoring step. Be careful here: if basically you find that your coverage starts to blow out when you put your tests, shift down to fourth, because you're probably writing too much speculative code you haven't got tests for.
If what I find though is that I'm saying, "I don't know how to go green, I'm not really sure how to implement this, and I want to feel my way with the aid of tests to probe me a bit more," then shift gears down and write some tests to help you do that. You could also use things like a REPL—REPLs in languages like Python, etc.—for this technique of just feeling your way a bit. But feel your way, but at the end of that process, **delete the tests!** The tests helped guide you, but now they're a burden to somebody else coming along. The person coming along who wants to basically change the implementation details probably means that your way of doing it, they're changing it altogether. Your tests are something they don't understand; we're just protecting your implementation details. Throw them away, and let the next person coming along write tests if they need to check the implementation, okay?
ATDD, very quickly:
Kent Beck originally said don't do it, because it has two problems: customers are not interested, and it's very expensive to do, right? He later, in 2011 or so, had a book about ATDD in his series: "Maybe I was wrong, maybe actually you need this because programmer tests are a problem, they don't focus on the gist of the system."
James Shore, who worked with Fit, who was one of the guys that wrote one of these tools, says, "I don't do it anymore." They said the reason, in fact, over the years, that customers essentially don't care, and they're a significant maintenance burden, right? So all these tools like Fit, Gherkin, that kind of thing, right? He just says actually, if I have customers again that help developers drive the developer tests, so the unit tests are informed by the developers talking to the customer and write the appropriate unit tests...
Over time, I used to believe in these tests; I used to think that was worth writing, you know, Fit and Gherkin tests, Cucumber, SpecFlow tests, because I thought it helped keep developers honest, we'd know what "done done" was. But over time, I've come around to realize that probably what we were solving was the problem of how we were doing our unit testing. Because our unit testing was focusing on methods on classes, we had nothing that told us about behaviors! But what we actually needed to do is shift our TDD practice to be about behaviors, and then we can drop ATDD altogether. Just write developer tests in the same language you're using to implement, and drop this horrible translation layer between natural language and your codebase.
Behavior-Driven Development:
I mentioned briefly Dan North, basically the guy who originated it, had this original post which basically describes testing to be about behavior. It is now a full-scale agile methodology; you may be interested in that. But this talk is not about Behavior-Driven Development as a methodology, right? It's taking that original kernel idea that Dan had of saying: test behaviors, don't test methods on classes.
Mocks, briefly:
Mocks are useful when a resource is expensive to create and represents a shared fixture problem. But the problem with mock objects essentially is people have used them to isolate classes. Don't do that! And especially don't do this thing where effectively you say, "I'm going to make sure junior developers have done the right thing by mocking. I force them to use mocks so I can see how they implemented the method." It just couples your test to your software, means you can't change it; it's never worth it. They over-specify software.
Generally speaking, IoC containers are overused. Try and use them as little as possible. Only use them because your framework demands that you actually do it because they don't understand how to create objects that you have. If your framework was good, it would actually ask you to give it factories, not an IoC container. And you might implement your framework factory in terms of an IoC container. If you find yourself implementing an IoC registration in your test, something has gone horribly wrong, okay? Bit like ORMs, we overused IoC containers; please avoid that, okay.
Right, quick summary:
- Reason to test is behavior, and basically because we have a new behavior.
- Write dirty code to get green, and refactor.
- No need to test refactored internals and privates; they're open season against tests.
- Write on a port in Hexagonal Architecture.
- Add integration tests only really to the extent you're covering basically how ports are implemented.
- Don't mock internals; ports are adapters, not one little bit missing out of defended code.
Okay, I ran out of time for questions, I'm sorry. I'm going to let you all go. Please come up to the front if you want to ask a question at the end, or come and find me during the conference. I'd love to talk about this stuff. You can see this talk online; versions have been given in places like NDC, so if you want to share with your colleagues you can go and dig out the video, and I think these guys are videoing it too, so you can spread the word with your colleagues. It's worth going and looking at that conversation between Kent Beck and DHH, mediated by Martin Fowler as well, for a similar conversation about how it's all gone wrong.
But thanks very much! Thank you. [Applause]