La histórica conferencia de 2009 de John Allspaw y Paul Hammond

My name is John allspaw. I run the operations group at Flickr. My name's Paul Hammond and I run the Engineering Group at Flickr. Today's talk, today's talk is going to be a bunch of different topics. Actually a vehicle. If you will to describe how development and operations fits together and gets along, actually work together and aren't huge assholes to each other at Flickr. But before we get started, we should talk a little bit about what Flickr is can ever remember? the hands, if they've heard of flicker, Okay, so for those of you who don't Flickr is a photo sharing website. At the moment, we store around three billion photos and any given second of the day, we serve around 40,000 of those photos per second. They take up about six petabytes of storage. You know what? A lot of kittens. There's a lot of yeah it's big. So yeah. So we're going to talk about Dev and operations historically traditionally. Even now this is usually thought of as Dev versus Ops High lurkers keynote had the had the graphic of, you know, two guys like this. You hear a lot of this. It's not my machines at your code and then you hit me saying. It's not my code, it's your machines. These this sort of behavior here and stereotype creates these sort of archetypes of developers and operations. Some developers might be construed as a little bit weird. They're really hope I in math land or and the Ops guys, man, they freak out every time. Goes wrong and they're easily excited. They may or may not drink too much. So what this what this in all seriousness, what this, what this leads to is The Stereotype and of go ahead, but the other stereotype we think of, when we think of operations people is, is of a grumpy old, man. Yeah, a grumpy old man. That says, no all the time. They're afraid that all these newfangled things are just going to break the site. Very, very her pointy full of blame. Yeah, so they say no all the time. Because the site breaks unexpectedly. Because no one tells them anything. Because they say no all the time, you know, you're now you now you and now that stereotypical ops manager, grumpy guy is the guy that says, no, I don't want to do that. Well, who wants to work with that guy? Nobody because the guy's a prick. So when we look at the roots of these kind of stereotypes, it comes back to traditional thinking about the role of developers in the role of operations. That the way most people think about it is that developers job is to add new features to the site and operations job is to keep the Stable and fast. That that's true. Right? Or is it I don't think that opposite job is to keep the site stable and fast. It's not their job. This might be a news flash to some people officers job is to enable the business, right? If this, if the business requires that the site go down every two weeks, even though you're the largest online gaming platform and you have millions of paying customers, Those paying customers might be quite fine for you to have availability of 97%, screw the five nines and this is just truth. It just so happens that keeping the site stable and fast is something that you see as common across business requirements. Speaking of business requirements, they're one of the most. I mean, one of the realities of working in a business particularly online businesses is that the business requires change if your business is standing still Then you're going to get taken over, you're going to be overtaken by some upstart like Twitter or Facebook jerks the problem you have, of course, is that change. If you look at the root cause of most outages and you generalize then, you come to conclusion. That change. Is that is that root cause of most outages? Most out of this wouldn't happen. If there hadn't been a change, a few days, a few hours, a few weeks previously. so, you get two options, you can discourage change in the interest of stability, I'm old no or you can be smart and you can build tools and culture to allow change to happen as often as it needs to. So that's what we're going to be talking about today. Most of what we're gonna be talking about is lowering the risk of change through through good, use of tools and good, working culture in your team. What we're trying to do with a lot of these tools is increase the Finish that that any given change isn't going to cause an outage or some kind of problem on the site. And then we're also going to look at ways that you can increase your ability to recover from those outages. Should they happen? of course, what really helps in this regard is when you have operations people who think like Developers, And developers. You think like, operations people? It's probably a bunch of you in the audience are thinking. That's me, I do that, I do that already. So we're going to do some handwriting and it's a bit hard to see because there's some really bright lights. But how many of you in the audience, consider yourself to be a developer? That's a lot of that, I guess. How many of you consider yourself to be on the operation side of the fence? Hmm, 2/3. And how many of you are in a job at the my where you're doing both? Special fill. So the operations people in the audience. How many of you have changed user-facing application code recently? The handful. How many of you work with developers? That are happy that you made those changes? Ha, there's something funny about that but the developers how many of you have worked out of hours because there's been some kind of sight problem and you've got phoned on an evening on a weekend. So that's maybe a third of the developers in the audience how many of you carry a pager? Or on or Uncle who are developers who developers and developers. How many of, you know, how many web servers you can lose before your side Falls over completely Israelis good questions. I think the answer is you can always think more like them. Yep. So we're going to talk a bit about tools now. And what we're hoping to do in this discussion of tools is these are some of the tools that work for us and they're not necessarily going to work for everyone and throughout we can ever give examples of the specific tools that we use. But the important thing that we're trying to get across is the concepts that are important and this is going to be a common theme throughout this conference. But if there's one tool that you put in place, if there's one kind of technique that you use automated Is that technique? It gives it just makes the operations jobs possible. If you've got more than a dozen servers, then manually managing every individual server is just not really an option. And from the developers point of view, it gives them a consistent predictable platform that they're building their app. On top of it means that they don't have 10 web servers running Apache, one point three and then three web servers running Apache, 2.2 unless you're in the middle of an upgrade in which case they'll know that that's happening. Yeah but without that consistent platform it's really hard for the developers to do their job. So we've got all, we've got all kinds of tools, this talk isn't about that Adam and Ezra are going to give a talk a little bit later about it and half a dozen better presentations on the topic. But the main point here is you have this, this this idea that you've got an OS image and you've got some sort of role that this server, this piece of infrastructure or cloudy bit actually. Lee actually performs, it's a task-driven infrastructure and really there's no difference. The difference between this and the cloud is that this has a blue cloud around it. And that's it. It's still the same idea. So, if you're, if you're running your servers on saec tooth and the OS Imaging side of things is taken care of by am eyes, but you'll still need the role of configuration management layer on top of that. Yeah, maybe even more importantly. The second tool we're going to talk about is Version Control, and I don't know many development teams that will even try and operate without Version Control. An increasingly, I'm seeing more and more operations teams, using Version Control, that something that something that in fact, we even used to do it used to be the case that Flickr source source code was in CVS. But all of the operations, all of the packages and all the configuration management was stored in perforce. And what that meant was that the developers had no idea? Was going on. I have no idea how to check out the perforce. Repository, John had no idea how to check out the CVS repository by having one shared revision control system. Then everyone on your team knows where to look to find the latest instance of the configuration for a particular box or they can see what's going on in the application where the changes are. This is really, really useful when you have an emergency last on Friday, I got find out, I was out having a meal and there was a problem with of our site and Kevin who works on John's team, phoned me up and because we have one single source code repository. Then I was able to talk through over the phone. What the single change that needed to be made to, to mitigate the problem that we had was, if we had different source code repositories that maybe Kevin wouldn't have access to it, you know, I'd have had to have gone home got up my laptop and make the fix myself and so this single Source control, it provides transparency and that's a very useful thing. From a development point of view, one of the most important things you can do is to set up one step build. And what we mean by that, is everything you need to do to take the code that is currently in your Source control system in SVN and turn that into a set of files that can be copied onto a production server and run the site. The screenshot that I'm showing on the screen right now is part of flickers internal development admin interface, this is the stage and build the build and Stage button. The button at the very bottom of the screen, the one that says form staging, you click that button, it performs an sdn, check out. It does all of the translations, that compiles, all the templates anywhere, we have compilation that we do for optimization, it does all of that. And then it copies that code onto a staging service that we can test it matically automatically, which means that you don't have people running this command and then you run this command. As it turns out, computers are really good at running commands. The same time and are The same order over and over again, you'll say, don't get this situation where a developer has done a build on their, on their workstation and that has subtly different configuration than developer to. And so subsequent changes, even though maybe there's no significant change in the code, the application behaves very differently when it's deployed. Once you've got that one step, build the next thing you need is a one-step deploy. This is flickers internal admin tool for deploying. At the top, we have a deploy log and that's a very very cheap form of change management. It means that everyone can see what's going on and can warn that maybe now isn't great time to deploy because there's some other change going on elsewhere in the system. And at the bottom we have a button. It says, I'm feeling lucky and you push that and it pushes the code out to the site. This is what it looks like. When you're pushing the code out to the site and the same principles apply here by making it one button, it means that there's very, very little room for error. It means that you're doing your builds and you're doing your deploys in a consistent environment. It means that there's no manual steps that might go wrong. And so, if you're looking at, if you're looking at the distant source code, then you can be fairly confident that that's going to be the only difference that you can see in the Performing the behavior of the application when it's deployed. And this is, this is the way, this is the way we do things, but you see this as a trend continuous deployment and continuous. Integration is starting to show up in a lot of operational tools, and, even vendor selling things, and, and open source projects as well. So, it's just a good idea. One of the greater things about this process is now, we have a deploy log. Let's just not skip over that for a second. We know who We know when and we know what? It's really, really important. John Adams, showed earlier they put, we'll talk about it later but they put the deploy timestamp at the top of their monitoring and metrics tools context is absolutely everything. And, you know, the title of the talk is. What 10-plus deploys had a, you know, you you can't Pretend or try to deploy 10 times a day. If you go down 10 times a day, that's not being agile. That's just being retarded. One of the points I want to get across here is that it's not just about having a one button on a web page that does the entire process, it could be that you have a make script, A make file in the root of your Source control when it comes to deploying, it could just be that you have a shell script that you run. It could be that you use Capistrano. It could be that you use. I don't know. RPMs and so, the deploys system that I just showed you is how the main flicker application is deployed. One of the things that It was starting to experiment with some of the smaller apps around the edges is use continuous integration server. We're using Hudson, to automatically generate Package packages, that then can be deployed, two boxes by the operations team. Again, it's a single step, build, you commit. The files for SVN and that you can push a single button in order to generate that package. And then and then a single step deploy, which is just someone from the Ops team will then actually deploy that package or a Dev. And yeah it's John's already mentioned because I build and deploy systems completely automated. We can do it more often and that allows us to make each deploys. Smaller to each individual deploys, introduces less risk. And should anything go wrong, it significantly easier to work out. What might have happened that so that we can recover The next thing I'm going to talk about is again, quite developer focus. And that is what we call feature Flags. You can also think about this as branching in code, when you think about revision control, branching systems there, a reaction to a reality of building desktop software, your Microsoft, you've developed Microsoft Word, you launch Microsoft Word, 1.0. And then your development team immediately starts working on Microsoft, Word, 1.1, and you launch from Ike's, you launch 1.1, and then a few days later, Realize this, a critical security vulnerability. So you have to go back to the 1.0 code base, and making change to it and ship that and then you release one point two, and then you have to make another change. And so, you're releasing three different versions of your software at once. For those of us that have a single instance of our web application, that's not the case. All that we have is what is currently deployed and any old versions, don't really matter anymore. The way we deal with that is that we always ship trunk. That's not to say that we do all of our development on trunk, although we do, you could do all, you could do your development in branches and then and then merge those branches. But by always shipping shrunk, it means that everyone on the team knows exactly where to look for the version of the code that is currently deployed on your site. It means that you don't have, you don't have to wonder, which, branch of which patch release is the one that's currently released. You just go into a sphere. And it's there, it's trunk. I mentioned that we don't do branching in SVN we do are branching in code instead and what that means is whenever we're developing a feature that we don't launch straight away, then then we block out those those particular code paths using conditionals. So here's an example, from PHP is an example from Smarty And that means that we can actually we have all of the code for every new feature. That's not yet released In production on our production service. It's just the configure set up so that so that it's not it's not actually visible yet. If you can't, if you can't tell where we're going with that, you'll see a little bit later, that smells a lot like an operational lever, talk about that later. We'll get to that in a minute. This allows us to do a couple of really neat tricks. The first is it allows us to do private betas on production server, on production Hardware with production traffic. Now obviously we have some really great The environments that are pretty realistic test and we do it. We do a lot of QA in those in those environments. But something we found out is that if you use a second set of servers for your testing, then you might notice, you might not notice changes in configuration between between your beta service, in your production servers, even with configuration management it happens. And we've had cases in the past where a new feature worked perfectly on a base servers and then the moment we push it to production. We realized that the backend internal web service that we were relying on was blocked from From our production dubs. And so, the feature didn't work we had to put out an emergency change request, so that back-end web service. It allows us to do bucket testing and this is incredibly useful. It means that we've got a new code path. If we've got a new back-end system that we're thinking about using, we can start to just push five percent of traffic to it, then increase it 10, 20, 30 50, we can turn on a feature for a subset of users, and we don't have to do any careful juggling with having different versions of the software running on different servers and pulling boxes in and out of production. We can just we can just do it all in code. And lastly, it allows us to dark launches, which is something Jonathan mentioned when he was talking about Facebook earlier, we launch new features behind the scenes. So if we've got a new feature that's going to cause more load on our memcache, boxes on a database servers on a search cluster, then then we just turn that feature on behind the scenes and start fetching the data for a few weeks, but not show that data. And then when it comes around to actually launch this new feature, we know we're going to be able to take the load because we've been taking the load for several weeks. It's a huge. And for four operations because there's no just takes the suspense. Sorry, fear out of the out of the gig in a huge way, right? It's not. It's it means that. Did you give the example of the new liquor homepage? I didn't know. Yeah, so I mean the we designed the new flicker home page. It's got a lot more information basically activity from your contacts and various amounts of activity, stream me Real Time stuff. It's a huge shitload of data, right? And so what are we going to do? Every pay, every home page. We're going to serve is going to grab a boatload more data from the databases will databases suck. That's a huge pain in the ass and and, and really scary. So we did it. We dark launched it for for a number of weeks. Yep. And the application just takes it and then Shucks the data. Well, it made the actual launch completely anti climactic, which is a really good thing. Yeah, should be. It should be obvious at this point. These feature Flags or disable Flags. How many do we have right now? Probably a couple hundred. We got a couple of hundred we can just turn things off, or we can change the behavior of things that that are affecting the site either availability or performance. If database, cluster happens to have some degradation or some other back-end service, has some degradation and its features specific. Then we can change this and without affecting the site, you know? And and do we roll back kind of not really a very often. What we do is we roll forward and just turn s* off. Some of those feature Flags also aren't just on or off. Some of them are, you know, like Paul said, some of them are knobs that have, you know, and variable variable amounts. So It's good. Shared metrics. Just like Version Control just like having shared Version Control. You can see my bits. I can see your bits. Well, we gather metrics right? We Gather shitloads of metrics, you know, here's We Gather all kinds of metrics right? Because the screenshot of our ganglion stall and we've got something like 37 different clusters. I don't know how many tens of thousands of metrics we grab on. Any on any given day at any given second. The important Point here is that devs not only know where this is. Where this is, they have access and they watch it. Just as obsessively as operations if you walk around the office, all the developers have at least in one tab on their browser, something like this. And and and and part of this because they know where it lives. It means that we can we can squirt application Level metrics in there, you know, CPU and network and memory and disk the all those things are, you know, all very important but they are really important when in the context of the application, right? I don't care if you're doing 75 percent, user CPU, how many Of that. What that role is that servers is doing? Widget eating. How many widgets is it eating per second for each piece of each each bid each percentage of CPU? This obviously comes into into capacity management and capacity planning, and all that sort of stuff if you think that's important. So, this is, this is a graph of This is oh yeah. This is great. This is a graph of the average over the last minute of a particular server that does image processing. So when you upload it, your kitten to Flickr, we cut it up into six or five or six different sizes and this gives you how long it took to process. Each one of those on average. And this is a graph of how many background tasks we have in our asynchronous task queue. And one of the interesting I actually made the metric to that, that is being shown in this graph. And one of the reasons I made this craft not because I was asked to buy John's team, I made this graph because I wanted to have some idea of how our offline task system was performing. The interesting thing here is, we've got this Dynamic where the developers want to make these traditionally. No metrics and John's team, make it easy for us to make these. We have a framework. Now we're, we're any part of the application will just write a file that has key value pairs in it? And our ganglia setup will just slurp it in and make pretty rodg graphs for us. This then leads to us creating adaptive feedback loops in the application. So where we have things that are happening asynchronously then we can have the application start to back off and try and put less load on the system if the system isn't performing well. So the offline tunnel syndrome that I just mentioned it will actually dial back. If the databases are starting to get overloaded to make sure that the database is already there to deal with the real-time load because if we need to run an offline, task 10 minutes later. Not a big deal and when we did the migration from Yahoo photos to Flickr. Then then we throttled that based on the amount of free disk space left on our storage to make sure that we didn't run out of storage. Yeah, that was a really good example. So we've got this multiple month process, right? They're shutting down Yahoo photos. One of the options to do with your Yahoo photos is you can go to Shutterfly or a couple other places or you could migrate them to flicker, right? So essentially what is it? It's Enormous massive exercise in. Cueing asynchronous tasks, it's going to take shitloads of time to get all of your 10 years of Yahoo photos imported into into flicker. In the same in the format that Flickr has with all of the metadata and it's going to take a long time so it's going to have to be yes, I'd like to move to flicker. Okay? We'll let you know. And and this is petabytes of data of petabytes of shitloads of database data, but petabytes of image data, Need to be processed to the different sizes because there were large differences between young voters and and Flicker. And so as far as you know we're pretty tight on and pretty dialed in with what how much storage is going to come online and when and but what we didn't know is that huge unknown variables. How many people are going to click that button? Say yes, let's go to Flickr. Well, that was really variable dependent you throughout the entire length of the, of the migration. So, Measuring when we're going to predicting, when we're going to run out of space and changing the and the migrations to adapt to that was a huge, huge win for us. John's kind of stolen or thunder on this one. It was our idea. Yeah. Adams Twitter and what I can give us credit, right? We do this. We actually said it already. We do put the time of the last side deploy up on every page of every metric we have. And that often makes it quite easy to work out. Why? A particular graph has just doubled or halved? This is an example where one of my team rolled out an optimization to our image. Code and we weren't sure what effect it would have. We thought it might be a little bit a little bit and it turned out to be quite a lot. Communications Starting Gate close towards the end of the tools bits. Here we use a lot of IRC at Flickr like a lot of other places and we use it for ongoing dialogue between developers and operations. What's going on with the site? What are you working on? We've got a lot of people with a lot of balls in the air and ircs helpful especially for remote people who work remotely. So we've got this conversation going on and it's good to have context so we do. Something Last.fm actually, wrote a great little tool to do something exactly like this. We squirt events computer driven events into the stream in IRC. So for example, certain alerts and monitors that we really care about build logs the deployed logs, when something was deployed. So, you know, you're having a conversation. I don't know. I'm having this problem with some rewrite rules and deploy and I don't know what do you about that are no alert. Hmm. So then what we do actually is we do we take all of this, this logged information and we shove it into a search engine. So it can say, what the hell did we do on Thursday, two months ago? What was going on? Do we see this problem before so really really helps because it means that humans now have context and not just you know, rrdtool back in time context there's actually human context which leads us to So all of these tools won't really help you if you install them. But you still have that incredibly argumentative combative culture going on and one of the other things that we think makes our working life so much easier is the culture that we have at Flickr just like automated infrastructure. It's like Ground Zero for sorry, like the very basic thing that you need to do in terms of tools when it comes to culture. The most important thing, the thing you need to start with is having a culture of respect. It's a bit hypocritical of us saying this, isn't it? But one of the most important things you can do, is to avoid stereotypes. I know that you worked with a cowboy five years ago, that just didn't care about your application on Justin care about your system uptime. But but not all of the developers that you're going to work with are as bad as him and not all of the operations people are going to be as obstructive as him. If you assume the worst in everyone, then everyone's going to assume the worst in you. Everyone. It's a delicate Little Snowflake. Everyone is a little bit of water. You every one of you, it's important to respect different, people's expertise, and different people's opinions because they have, they have different experience and they're going to come up with different solutions to problems than you. And those Solutions might not be the best solution, but you should at least respect their their, their suggestions. And the other thing that's really important is to respect different people's responsibilities, which we all have different responsibilities. Business. John is the one that's going to get a hold in front of management. If Flickr. Spit has too much down time. I'm going to get hold in front of management if we don't launch a new features in the next two years. And and so that means that we are going to have different different priorities and it's important to understand recognize and like I said, we respect those. Part of that is, is when you have conversations about problems, you know, just saying, no is another way of saying. I don't care about your problem, you can't write a bloom filter. So screw you finding out what trying to, what problem are they trying to solve with? It's a developer trying to fix a problem or an operations person trying to fix a problem. This is where you find the Most cool stuff, you know, almost all of the successful companies that have that have reached elbows of scale and hurdled over them is where is where developers and operations come together to come up with unique Solutions. You know, trying to think of examples, memcache is a perfect example of this, right? It used to be that databases were like dbas and Ops people. That's there. You know, those That was them. I don't know. I'm a developer going to write code and there's a magical day to be somewhere that does answers for me, you know, memcache was written and various architectures were designed to solve this problem. And the only way that it happened is when developers and operations work together, If you're dealing with someone who is just saying, know if you're trying to solve a problem and you know, that the response from your colleagues in the other team is going to be to say that that's a really bad idea. Then hiding your solution from them, it's just a really, really bad idea. If John's going to say no to something that I want to do, then it's probably for a really good reason. And by hiding it from him, I'm not giving him an opportunity to provide his expertise. More importantly, if I'm gonna hide something from him, he's going to find Out about it eventually, and he's going to be pissed when he does. Again, for the developers, one of the things that you can do to really kind of get that sense of respect about your work, is to talk to Ops upfront before launching something up for proposing. Something about what the impact of it's going to be, if you're going to push some code live, what metrics are going to change? What what boxes are we going to see more CPU usage? What boxes are going to get more free memory? What are the risks? That something might go wrong? What are the signs that something is actually going wrong? So what should the operations team be looking out for? And if something does start to happen, What are the Contingencies, how can Ops recover to make sure that the site carries on working? The problem with coming up with these answers, is, you're going to need to work out these answers before, going to talk to Ops, and you might not have all of the answers. But you should at least use this as the base of a basis of a conversation. Which leads us to trust that last slide has a bunch of things and sort of recommendations to help frame this conversation. It means you know if a developer comes to Ops guy, grumpy Ops guy but he's not going to be grumpy for too long and says, hey, I'm thinking about pushing this. I think it's going to change low characteristic on cluster, a, b and c. I built this, there's this little hook. You can set this to 0 if something goes awfully wrong. Blame me for it. We got to do it. Big enough to put this piece of code in because of this feature, blah blah blah. Now I'm off person is like s. This guy was really thinking and not only that, but he cares about this. You cares about the uptime of sight and he cares about not waking up my team in the middle of the night. Which is about, trust Ops need to trust Dev to involve them on featured discussions. So you're talking amongst yourselves about. We got a feature new awesome and you know think to yourself, hey, should I should I talk to Ops about the maybe we should tell Ops about this? And I'm not talking about the, you know, I'm not talking about regular operation stuff. That should happen early, like capacity planning that should that's a given. I'm talking about changes, that happened to the site. And the development team needs to trust. The operations are going to discuss infrastructure changes with them ahead of time. They're not going to find out that the PHP has just been upgraded to php5 the day after it happened again. There's it kind of, it sounds like a really obvious thing to do but I've worked in some really dysfunctional teams in the past where this wasn't accepted as a necessity and it comes back to everyone trusting that everyone else is going to try. The best for the organization as a whole. Yeah. Yeah. If if you know if you think you know, I don't want to tell Jim because he's gonna freak out so we're just going to do it anyway. Well that does that just means that you're a f cowboy and and and you don't want the office people doing that either and Cowboys are losers. So Something that something practical that we do to help. Make sure this conversation happens is as much as possible. We try and create shared but run books and shared escalation plans. So, so John and I will sit down together looking at a new feature or members of our team, will sit down together and work out what how exactly this new feature is going to be supported operationally and what what the scenarios are in, which it might go wrong, who needs to be involved in fixing those and and that just forces Conversation to happen around. What are the risks? What are the contingencies? And it made sure that everyone's on board. Cancer. We can't say this too much, providing knobs and levers and knobs and levers for developers to Monitor and and watch the code that they've been slaving over and and developers writing Hooks, and knobs and levers, such that these things can be operationalized. It's not enough for you to just throw a piece of code over the fence, but has all of the assumptions around. All of the variables pre-compiled in such the operations, can't change it. If you think that having 20 Child processes is good, it's a good number. You need to make that something that Ops can configure so that if they upgrade this Hardware underneath it, then they can, they can run that themselves. Controversial Sly, do you want to take sure? I believe that all people should make sure developers can see what's happening on this system. Without going through operations, there's nothing worse than having to play phone tag with shell commands, just dumb. By the way, developers are writing the code that are running on your machines, and if you have a big, if there's, if you are, you know, of course, guardrails are one thing like like, highly guarded said in his, in his Keynote, But, giving somebody a read-only shell account, even if you're super paranoid on production Hardware, it's low really low risk, and they can actually see what's happening, you know, beyond the metrics that you're now sharing because you're sharing type of operations person. It means that they can actually get into the guts. Make sure that they can access all of your all of the metrics, not just the ones that you haven't ganglia or nagios and that sort of thing, they should take a look at at what's going on and you Shouldn't be afraid to allow them to do it. We're not saying that every developer should have root on every production box, otherwise called root on some of them. But but the read-only account is a very, very low risk. And and as a developer, it's very, very difficult to diagnose things that you could be doing better in your application. If you can't even see the files, the lock files that it's making and you can't even see the the process tree so you can see how much CPU it that individual. Patients. Taking up is very, very difficult to help operations if you don't have access to machines. Okay, the third, the third aspect of culture that we want to talk about is what we think of, as a healthy attitude around failure, the important thing to realize here is that failure is going to happen. It's not a question of. If it's a question of when and if you're spending all of your time thinking about how you can prevent failure, then then you're not spending any time considering how you're going to respond to failure when it happens as an example Airlines, they spend, you know, Airline Spend, they spend hours days each month in simulators planning for what will happen. If you'll cuts out. What happens if they lose an engine, what happens if, but then still fight for land on the Hudson? Yeah. But still, they also develop procedures to help them when something goes wrong and and, you know, evacuation plans and so on. If you want to think about this one way to think about this is if you're in a site outage, you know, if you're having some kind of medical problem, would you rather be treated by an EMT that deals with a heart attack once a year? Or would you rather be dealt with by an EMT that deals with a heart attack? You know, every few weeks. So one of the things that you need to develop in both your operations team and your development team is the ability to respond to problems quickly and effectively, of course if you're dealing I'll choose every single week than you get bigger. Problems got bigger problems so we come back to the idea of fire drills. Whenever we have an instant flicker, then we'll have whoever's on call from the operations team. And maybe a couple of senior Engineers working to resolve that problem as quickly as possible. One of the things that I've started doing is sending a message to one or two. The junior Engineers on the team saying, look just assumed that no one else is in the office assume sites just gone down. You're the only person around. See if you can work out what you do to fix it, don't actually make any live changes, but just see if you can do diagnose it at the same time. And after the outage is over, then then we compare notes and that's a way of just training. Some of our more, junior members of the team how to respond to these kind of problems. Avoiding blame. It sounds like a really, really basic thing. We've got a rule of no finger pointing it's not a rule that we really have to enforce actually because we have this, I'm extremely lucky to be involved in an organization that will actually take blame as soon as it comes really close. Let's go, let's go. So yeah. So here's like a regular problem problem in a nun. Evolved Old. So the problem happens and then oh my God, what the f, what am I going to do? Maybe it was gonna change. I'm not sure. I don't want people to know what's going on. I'm going to delete the logs, what's going on and, you know, chicken head cut off and then it's not me. It's not my code, it's your machines. Then you finally get around to figuring it out. There's a shitload of waste time, right? You've got territorialism getting in the way of fixing stuff. Here's here's a Question, you could you could do it this way, you could figure it out. You could fix things and then later, if you want to feel guilty about it, then you can. This is generally what happens to, this is generally, what happens if Flicker? And actually, if anything both in my team and in Paul's team, we end up having people just falling over themselves, trying to prove actually that it was their fault, that broke it because, you know, ten times a day, who knows, right? But then they can be the guy that fixes it. And they can be the guy that fixes it. Something weird going on there. Yeah. Yeah. Again underscores something where he said, remember. Developers, when you write code, somebody else will wake up in the middle of the night. There are those organizations that put devs on call? If you think that something might happen? The middle of night and you're not going to be around or you're going to be at Thanksgiving and saint-martin other people are going to fix your s. Even if even if even if you have dibs on call is generally the Ops Team to get this is Jerry first page. So and if you're if you're in the scenario where you've broken something in the middle of the night and you come in the morning after and find that, that it was your fault, that there was a half-hour outage and that five people got paged and woken up. Then you can just say sorry, makes people feel better about it. Yeah. And if you don't say you're sorry it, you know, it sends the message of well, screw you. Aren't you getting paid extra for being on call? Anyway, and something I see a lot of velopment teams do is they rely on the fact that operations are there? That operations are going to get woken up in the middle of the night, to fix their stuff for them. One of the questions you can ask yourself is what would I do differently? If I didn't have someone there picking up my slack and if you can think of things that you change, if you were the one that was getting woken up in the middle of night, then maybe you should be changing those things. Operations should provide constructive feedback, continual feedback on how things are going, you know, these guys are writing code, to make your business run, and to get more users and to take over the planet. So they should know how things are. Don't don't, don't gripe about it, explain to them what's going on and say, hey look you know, I noticed that every, you know, that this Cron job that happens every 6 a.m. but blah, it's not that big of a deal. Oh, but it does go, you know, explain to them. There's you know there's all kinds of metrics. You know what you want to look at. They know what they want to look at, just share the information. So let's summarize the six tools. You should think about using and four aspects of your culture. You might want to think about changing automated infrastructure, shared Version Control One, Step build and deploy feature Flags shed, metrics, IRC Bots, respect trust healthy, attitude, towards failure, and avoiding finger-pointing. To be clear. This is definitely not easy but you're you're free to continue on freaking out on each other if you want. Thank you for your time. Thank you.

  • a2/3/allspaw-hammond.txt
  • Last modified: 2021/08/27 14:30
  • by jherrero