<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
<title type="text">dbsmasher corner</title>
<generator uri="https://github.com/mojombo/jekyll">Jekyll</generator>
<link rel="self" type="application/atom+xml" href="https://blog.dbsmasher.com/feed.xml" />
<link rel="alternate" type="text/html" href="https://blog.dbsmasher.com" />
<updated>2026-07-02T23:32:06+00:00</updated>
<id>https://blog.dbsmasher.com/</id>
<author>
  <name></name>
  <uri>https://blog.dbsmasher.com/</uri>
  
</author>


<entry>
  <title type="html"><![CDATA[Not My Job]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2022/05/24/not-my-job.html" />
  <id>https://blog.dbsmasher.com/2022/05/24/not-my-job</id>
  <published>2022-05-24T00:00:00+00:00</published>
  <updated>2022-05-24T00:00:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;h2 id=&quot;it-is-not-my-job&quot;&gt;It is not my job&lt;/h2&gt;

&lt;p&gt;The job of a Staff+ IC (individual contributor) leader is murky at best. Even when I spoke about this in [the first event dedicated to Staff+ folks], I felt compelled to explain that “it is a job that can be different things at different times”. These differences can come from various sources:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;The stage the company is at when it hires/promotes Staff+ folks&lt;/li&gt;
  &lt;li&gt;If you are the first or one of many&lt;/li&gt;
  &lt;li&gt;It can be simply that there has not been time to put an effort in properly defining the role yet.
Having said that, there &lt;em&gt;are&lt;/em&gt; patterns that, I think, are worth flagging as inherently toxic for most staff+ folks and should be recognized as the outlier (and when they are narrowly ok) vs “this is just how the role is”. The last thing we want is for this role to become a burnout factory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This conversation started when someone in a people leadership position tweeted &lt;a href=&quot;https://twitter.com/thiagoghisi/status/1525073311369252865?s=21&amp;amp;t=BJtmKeShTO5OMnxp4aee7Q&quot;&gt;this&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/tweet_screenshot.png&quot; alt=&quot;tweet screenshot&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Now, I am sure they have their experiences that make them believe this to be the right advice to give. But I would like to offer a counter to this argument and why I think it is a road to frustration, ineffectiveness and ultimately burnout.&lt;/p&gt;

&lt;h3 id=&quot;the-difference-between-being-glue-and-filling-gaps&quot;&gt;The difference between being glue and filling gaps&lt;/h3&gt;
&lt;p&gt;Before I go into why I think this is not good blanket advice, I’d like to address the difference between what this tweet says and what &lt;a href=&quot;https://twitter.com/whereistanya&quot;&gt;Tanya Reilly&lt;/a&gt; has coined in her popular blog post and conference talk &lt;a href=&quot;https://noidea.dog/glue&quot;&gt;“Glue work”&lt;/a&gt;. Glue work is an important aspect of building software across multiple teams, making sure everyone is rowing at the same speed and in the same direction. It can be a great opportunity for building leadership skills as you grow from only focusing on a technical backlog to more abstract concepts like “is what we are building matching our strategy?”. And as Tanya says, it is essential when you are senior and potentially limiting when you are not. I would go a step further and say that a good career ladder for a growing engineering organization should recognize that glue work. It should recognize it as part of how to measure the performance of its senior team members and not let it be invisible work.&lt;/p&gt;

&lt;h3 id=&quot;glue-work-vs-gap-filling&quot;&gt;Glue work vs “gap filling”&lt;/h3&gt;
&lt;p&gt;I liken the work that goes under “glue work” to spackling paste that ensures that all parts of the software delivery for the business is happening smoothly.&lt;/p&gt;

&lt;p&gt;However, that is not the same as saying “as a senior leader, everything is your job”. Gaps can happen in many shapes and sizes, and you need to recognize when a gap needs more senior leadership attention instead of trying to absorb them all. In those cases, you should be working on properly communicating the gap and its risk to the business (and risk to &lt;em&gt;which part&lt;/em&gt; of the business) and NOT attempting to solve everything. Here are some examples of possible gaps that you should be _definitely _ saying “this is not my job” and instead communicating risk about:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;A team is severely understaffed, and their manager is not advocating for them&lt;/li&gt;
  &lt;li&gt;You are involved in planning with a part of the organization and can see clear signs of a toxic work environment&lt;/li&gt;
  &lt;li&gt;A critical part of the technical stack is a recurring source of incidents and no one officially ‘owns’ it in engineering.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You might exclaim, “but I can’t stay quiet!”, and I am not suggesting you pretend you saw nothing. But be cognizant of what you can and cannot control as someone who is in leadership without explicit authority over teams. You owe it to the engineering managers of managers you work with to communicate the gap, but you should not put it on yourself to try to &lt;em&gt;resolve&lt;/em&gt; the problem without communication.&lt;/p&gt;

&lt;h3 id=&quot;why-this-is-even-more-dangerous-for-non-manager-leaders&quot;&gt;Why this is even more dangerous for non manager leaders&lt;/h3&gt;
&lt;p&gt;This dynamic of “fill every gap” is specifically dangerous when you are a leader without an explicit scope of authority. When you are in a leadership position without people management authority over teams, resources to deploy to make changes you deem necessary, you run the risk of burning out trying to make changes that require authority. Authority you do not have. As we all learned in airplane safety videos, you have to put your mask on first. Taking on every gap and attempting to stretch to fix them all can become a crutch for poorly defined staff+ roles.&lt;/p&gt;

&lt;h3 id=&quot;why-one-is-healthy-and-needed-and-the-other-is-not&quot;&gt;Why one is healthy and needed and the other is not&lt;/h3&gt;
&lt;p&gt;It may seem that the difference between the two is nuanced and likely also contextual to the org you are in. And I would agree! Which is why trying to encapsulate it in a tweet is glossing over a ton of critical nuance that can mean the difference between a highly fulfilling job and one that leaves burnout and trauma.&lt;/p&gt;

&lt;h3 id=&quot;your-job-is-to-make-the-org-successfulwith-fine-print&quot;&gt;“Your job is to make the org successful”…..with fine print&lt;/h3&gt;
&lt;p&gt;The original tweet also claims that “your job is to make the org successful” and I would like to point out what a dangerous sentiment that is if you are a staff+ IC in a large, sprawling, org. Carrying the burden of “making the org successful” is what executives are paid to do. Your job, while strategic and larger scoped and more nuanced, still needs far more concrete boundaries than a very non-specific “fix it”. And it needs to come with clear mechanisms to help make you successful within the scope of things you &lt;em&gt;can&lt;/em&gt; affect. 
Let me say this part bluntly, hiring staff+ is becoming trendy right now. And like all past trends in recruiting, some companies have a good idea what it means to have highly senior/experienced ICs, some just happened to grow some accidentally and some think they &lt;em&gt;desperately&lt;/em&gt; require some without a clear idea of why or how to define success for those new levels.&lt;/p&gt;

&lt;h3 id=&quot;what-do-i-do&quot;&gt;What do I do?&lt;/h3&gt;
&lt;p&gt;So what do you do if you are in a staff+ role, and you suspect this is what is happening? Start by assessing all the projects/initiatives/conversations you are involved in and having regular conversations about. Do all of them align with longer-term goals you have with your manager? Does your manager even know about all the gaps you are ‘filling in’? (here is a major red flag for you). Assessing whether what you are doing day to day needs to be an intentional process, something you and your manager re-assess routinely and compare to your goals and the organization goals. Be very aware of being pulled into projects with no measurable milestones. Beware of signs of inaction if you escalate a gap you are not responsible for (such as poor product planning or severe understaffing).&lt;/p&gt;

&lt;p&gt;You should use your experience and influence to shed a light on gaps and risks. But you cannot fix them all.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2022/05/24/not-my-job.html&quot;&gt;Not My Job&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on May 24, 2022.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[High Performance MySQL]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2020/04/08/high-performance-mysql.html" />
  <id>https://blog.dbsmasher.com/2020/04/08/high-performance-mysql</id>
  <published>2020-04-08T00:00:00+00:00</published>
  <updated>2020-04-08T00:00:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;There are things that one never expects to get to do. I have talked about this on Twitter before but I still remember being a new immigrant who took a chance and quit college to come to the US, start over and merely dreamed of working tech.&lt;/p&gt;

&lt;p&gt;I have also talked about how I started in tech after college, how I never planned to be a DBA and it was all a ‘happy’ accident and a necessity of the time that showed me something I didn’t know I was passionate about doing well. Back then, the de-facto book on how to run MySQL in a production setting besides Oracle’s own docs was a book called “High Performance MySQL, 2nd edition”. MySQL 5.1 was brand new. InnoDB was becoming a more mature engine, people were starting to look at MySQL as a much more serious datastore technology to use in companies looking to rely on open source software to build their business.&lt;/p&gt;

&lt;p&gt;That book helped teach me so many things back then. What is a buffer pool? How does one optimize it? How does a query run? How do I understand a query plan? What datatypes cause more disk activity and how do I optimize all these buffers?&lt;/p&gt;

&lt;p&gt;Since then there has been a third edition that added lots of new content about 5.5 (which was then the ‘GA’ version).&lt;/p&gt;

&lt;p&gt;The third edition came out in 2012 and since then, there has been 3 major releases of MySQL and so many improvements, new frameworks and tools in the MySQL community. It is time for a 4th edition that covers not just new core features in MySQL but a new look at &lt;em&gt;how&lt;/em&gt; to run MySQL in a high performance, high demand environment. Not just as islands of database instances, but as a database platform that serves the business and the delivery teams. How to manage compliance needs. How to manage schema changes. How to scale not just the database instances but also how to choose the right access layer for your application layer. When to recognize that you need to go polyglot and branch out of relational databases altogether.&lt;/p&gt;

&lt;p&gt;I will be working on creating the new version of the book that got me hooked on scaling databases all those years ago. Planned for release in the fall of next year. I hope it does for many what the second edition did for me years ago. This book was the start of an amazing journey for me and I am thrilled to help bring the next edition of it to up and coming database engineers.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2020/04/08/high-performance-mysql.html&quot;&gt;High Performance MySQL&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on April 08, 2020.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[How We Keep Learning]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2019/08/08/how-we-keep-learning.html" />
  <id>https://blog.dbsmasher.com/2019/08/08/how-we-keep-learning</id>
  <published>2019-08-08T00:00:00+00:00</published>
  <updated>2019-08-08T00:00:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;h3 id=&quot;tech-10-yrs-ago&quot;&gt;Tech, 10 yrs ago&lt;/h3&gt;
&lt;p&gt;Back when I started in tech, the most common deployments were what is known as a LAMP stack. This was a kind of architecture that was simple to reason about. Also, we commonly deployed such a stack by either leasing servers from vendors like Softlayer and Rackspace, or by directly leasing space in datacenters and dealing with all the details of networking, cage space, cooling, power, etc. 
That environment created a generation of operations engineers who knew far too much about how the internet really works. Baling wire and glue, we liked to call it. You can just read one BGP incident review writeup to see how fragile it all really is.  And because that stack was simple and had ‘A’ database, it meant we could get away with deploying that by hand. Configuration would be artisanal to specifically match our work load, snowflakes and highlanders all the way down.&lt;/p&gt;

&lt;h3 id=&quot;cloud-allthethings&quot;&gt;Cloud :allthethings:&lt;/h3&gt;
&lt;p&gt;Fast forward more than ten years later, and a diminishing number of companies run their tech that way. For shops that are running at moderate scale or new startups that are still iterating and haven’t hit their stride yet, the common path is in the cloud. Not just through leasing EC2 instances from AWS but by using managed services that shorten the time to features for engineers. Need a database? Here is one that is not just up and running in minutes, but comes with write failover, replicas in multiple regions and basic monitoring and metrics all baked in. 
Need container orchestration? Here is a hosted kube cluster where all you have to do is provide the helm charts of what your deployments look like, its care and feeding and management is all solved for you. 
Need to deploy services? You don’t need to size some hosted virtual machine or manage the reservations of these instances anymore.  Use lambda to directly to deploy your code, set your concurrences and conditions for running and you are off to the races.&lt;/p&gt;

&lt;p&gt;We don’t need to learn how these complex stacks work under the hood, it’s all managed right?&lt;/p&gt;

&lt;h3 id=&quot;cloud-reality-check&quot;&gt;Cloud reality check&lt;/h3&gt;
&lt;p&gt;In reality and put bluntly, there is no free lunch. Not forever at least. Yes, managed databases and Kafkas and kube will be convenient and much simpler to use when things are new and you don’t have a lot of customers yet and you just want to focus on shipping new features but if you are truly banking on long term success and a growing market share, you cannot rely on managed services forever.&lt;/p&gt;

&lt;p&gt;All systems have limits and when you are using a managed service, it means you are at best going to get the uptime and reliability guarantees that cloud provider can give. Sometimes those guarantees are not enough. And even when they are enough, sometimes vendors don’t meet these guarantees and the best they can do is refund you some money or give you some ‘cloud credits’. And if you are already past the ‘growing pains’ stage as a company and have a large roster of very big customers who have uptime expectations of your service, you can’t just tell them “well our cloud provider apologized”. As Jeff Hodges says in one of the best websites out there: &lt;a href=&quot;https://whoownsyouravailability.com&quot;&gt;Who owns your availability?&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And this goes beyond just uptime nines and availability. As your product grows, you will learn the limits, sharp edges and bear traps in all sorts of managed cloud services. Aurora doesn’t like getting past a certain write throughput rate? Dynamo can have hot shards? Certain parts of EKS are not yet as configurable as you need them to be? These are all real things that can cause scaling issues, unexpected behavior in your software or at the extreme, outages. And the managed services technically did not have an ‘incident’ but like all complex distributed systems, they will have limits and the risk of your success is being the first, or one of the first, to find these limits.&lt;/p&gt;

&lt;h3 id=&quot;working-with-younger-engineerswhere-senior-engineers-come-from&quot;&gt;Working with younger engineers/Where senior engineers come from&lt;/h3&gt;
&lt;p&gt;Engineers early in their careers find the advice to ‘learn the sharp edges of your tools’ constricting them from providing value, from building things, from getting stuff done. And this tension can be frustrating to seasoned engineers who may feel the new team members are about to make the same mistakes they did in the past. 
That is not a surprising state of affairs. 
But it is also something organizations need to get ahead of and treat as an opportunity for learning. Things like design reviews, pairing between senior team members and newcomers, brown bags that tell stories of past mistakes. These are all things that help cross pollinate knowledge and level up the new team members. Or more importantly, fill gaps in the experience senior engineers you hire in. Because like it or not, even engineers with years of experience may have not encountered the things you had to scale against in your specific organization/product.&lt;/p&gt;

&lt;h3 id=&quot;learning-from-incidents-and-near-incidents-nurturing-inquisitiveness&quot;&gt;Learning from incidents and near incidents. Nurturing inquisitiveness&lt;/h3&gt;
&lt;p&gt;This is a topic that can take up not just a whole other blog post but entire books and academic papers and it does. But I would be remiss if I did not also mention in this post the huge importance of learning, as an organization, how to learn from both incidents and near-misses. This is more than just “how do we steer/facilitate retrospectives” or “what artifacts should we produce from such meetings” which are both important questions. But in the intersection of “learning from incidents” and “how we grow senior engineers” I am interested in how these learning exercises, when done right, can become a critical tool in transferring that invaluable ‘smell test’ or ‘hunch’ that valuable senior engineers have and showing less experienced team members how to troubleshoot these complex systems we have built. 
Whether or not you run in a data center and have to manage your own kube clusters or are glueing together managed services in a public cloud, it is the incidents and near misses and how you reason about them that evolve the team’s ability to know what their code &lt;em&gt;really&lt;/em&gt; does. Once that value proposition sinks in, it becomes clear how important it is that your engineering organization fosters an atmosphere of inquisitiveness.&lt;/p&gt;

&lt;h3 id=&quot;finally&quot;&gt;Finally&lt;/h3&gt;
&lt;p&gt;Invest in your engineers’ learning. There is no way around the need for this to be an explicit investment by fast growing companies. The cloud is for sure convenient. It is completely understandable why companies prefer to start new products there and orient rewrites of old things to move there. Engineering time is the most expensive asset for a tech company and all companies are now tech companies so anything that speeds getting new features in customer hands is a win. This makes it all the more important for companies that want to build a scaling engineering team to intentionally invest in learning. In the past this learning was happening ‘by accident’. Your old school system admins learned the sharp edges of tech by bleeding on them but now the cloud is making younger engineers less aware of the things that will cause issues at 10x (or 100x) scale. 
And for those of us who know the sharp edges exist, keep reminding yourself what it is like to have been a novice. It will keep you humble and make you a better mentor for the newcomers.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2019/08/08/how-we-keep-learning.html&quot;&gt;How We Keep Learning&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on August 08, 2019.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[On Being A Principal Engineer]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2019/01/28/on-being-a-principal-engineer.html" />
  <id>https://blog.dbsmasher.com/2019/01/28/on-being-a-principal-engineer</id>
  <published>2019-01-28T00:00:00+00:00</published>
  <updated>2019-01-28T00:00:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;h3 id=&quot;the-managers-path&quot;&gt;The manager’s path&lt;/h3&gt;
&lt;p&gt;I bought and read &lt;a href=&quot;https://www.amazon.com/dp/1491973897/ref=cm_sw_r_cp_tai_Aq9tCbSNCEHAA&quot;&gt;The Manager’s Path&lt;/a&gt; by the awesome &lt;a href=&quot;https://twitter.com/skamille&quot;&gt;Camille Fournier&lt;/a&gt; when it first came out. At the time, I was a senior Database Engineer aspiring to become a principal engineer in my organization. At SendGrid, principal engineer and principal engineer 2 are manager and director level roles respectively without having human direct reports. We had just overhauled the career ladder to provide a full technical ladder, and I now had the opportunity to grow that did not require going into people management. It is not that I disliked managers or that I think it is easy work. To the contrary, I very much appreciate the complexity of engineering management work but I felt that I wanted to strengthen my technical expertise and solidify my career as an individual contributor before considering going fully into people management.&lt;/p&gt;

&lt;p&gt;Now, 2 years later, I have been a principal engineer for a year. I am now closer to more technical strategy decisions than I used to be and I decided to read the book again. The title of the book and its stated goal might be laying out the path in engineering management up to senior leadership, but reading it made it clear to me that the book is also of great value to individual contributors to help explain what all these titles mean, what each layer of management is supposed to focus on, and how engineering concerns converge ultimately with business concerns and crystallize into a strategy for a technical organization. I also realized that while I am still an individual contributor, the principal engineer role carries enough cross-organization work, and enough people skills, that it is much closer to management than it may seem without engineers reporting directly to me.&lt;/p&gt;

&lt;h3 id=&quot;career-ladders-and-the-myth-of-a-flat-org&quot;&gt;Career ladders and the myth of a flat org&lt;/h3&gt;
&lt;p&gt;It surprises me that many shops still claim to have a ‘flat org’ or claim that they do not believe in titles. I have heard it said about data stores that ‘all databases have schemas…even the ones that say they do not’ and I think the same applies to organizations that are larger than a small handful of individuals. They may claim or even believe that they are not encumbered by the politics of titles and organizational hierarchy but that simply means that the power structure is there and implied and not based on clear milestones or competencies either.
When companies’ engineering teams grow past a handful, the engineering leadership has to document explicitly what they consider are the competencies of a senior engineer, not leave that up to interpretation. Implied competencies are easily colored by personal bias, both implicit and explicit. And these biases are a quick way to lose competent engineers.&lt;/p&gt;

&lt;h3 id=&quot;principal-not-senior-senior&quot;&gt;Principal. Not senior senior&lt;/h3&gt;
&lt;p&gt;One of the most common steps in defining titles when a company is hitting its growth stage is adding ‘Senior’ to the moniker of engineer but soon enough, especially if retention is good and people are staying on board a number of years, companies realize they need more than just a 2 level engineering ladder. This is where titles like ‘principal’ or ‘staff’ engineer become part of the defined career ladder. However, principal engineer should not be seen as a natural progression to senior engineering levels. It is an IC position that is on a different playing field as it involves competencies that no longer apply to only technical prowess. For an engineer to get to the Principal Engineer level, there needs to be cross organizational collaborative signals, there needs to be a clear understanding of architecture and design decisions that go far beyond the immediate technical area of expertise.&lt;/p&gt;

&lt;h3 id=&quot;a-force-multiplier&quot;&gt;A force multiplier&lt;/h3&gt;
&lt;p&gt;Once I became a principal engineer, it quickly became clear to me that my job involved a lot more than closing tickets and writing code to achieve things. Yes I am still a maker not a manager and I do not have anyone reporting to me but my new position now requires leadership duties that are best not left implicit or not handled with the same outcomes oriented focus as my past coding assignments. One of the most important competencies of a principal engineer is to become a force multiplier. This is a much more mature definition of ‘10x engineer’ than the Silicon Valley cargo cult likes to use. A PE does not produce 10x the features or fix 10x the tech debt tickets. A truly valuable PE makes their whole team better by advocating for best practices, gently reminding people of why the processes we have exist, and helping the less experienced engineers find ways to ‘level up’. A good PE can speak to technical aspects of the product, connect planned work to business strategy and to what makes the company more successful and maybe most importantly, have the interpersonal skills to influence others around them towards these goals.&lt;/p&gt;

&lt;p&gt;This is why promoting any engineer to PE without clear skills in more than just ‘the code’ would be a disservice to the team and a bad signal for the rest of the organization as to what things management actually values when it comes to the non manager career track. If you want to watch and see how the ‘brilliant jerk’ anecdote came to be, it is those promotions to senior IC titles based solely on code output and not all the value that inter human skills can bring into getting large numbers of humans rowing in the same direction.&lt;/p&gt;

&lt;h3 id=&quot;cheerleader-recruiter&quot;&gt;Cheerleader, recruiter&lt;/h3&gt;
&lt;p&gt;One thing the book especially focuses on explaining is “what does that person do, anyway?”. It lays that out by showing when/at what point a ‘manager’ changes focus from day to day tactical to the long term strategic. And few things are more strategic to a company than its ability to recruit great engineering talent.&lt;/p&gt;

&lt;p&gt;Managers carry a responsibility towards recruiting and representing the company well. That’s an obvious fact. As a principal engineer, recruitment is one of your responsibilities.&lt;/p&gt;

&lt;p&gt;When working at the senior engineer level, the focus is more on “getting tickets/projects done with little to no direction” and “being proactive in fixing tech debt or helping solve problems/bugs”. But once i became a principal, it became clear that by virtue of being one of a few, i now carried a larger impact on morale, organization culture and even on recruiting and representing my engineering organization outside of the company. Behaviours I display, either at the office or at tech events, are now a display of the behaviours my company rewards. My behavior is a signal to anyone who may consider working in my organization of whether we align in values or not.&lt;/p&gt;

&lt;h3 id=&quot;manager&quot;&gt;Manager?&lt;/h3&gt;
&lt;p&gt;A few years ago I was where a lot of senior engineers (&lt;a href=&quot;https://phys.org/news/2017-06-female-managerial-roles-unintended-consequences.html&quot;&gt;maybe more the women than the men&lt;/a&gt;) tend to be:. at a fork in the road. Do I become a people manager or do I stay in a technical IC role and focus on solving technical problems?&lt;/p&gt;

&lt;p&gt;Trick question! Once at a certain level, all problems are solved by people. There is no such thing as ‘purely technical problems’. In fact, this is the level where one wishes more problems were purely code because we can make code do a lot of things. Making people do anything is harder and influencing people to do what we want is harder still. This is what principal engineering is about. The role is far less about solving intricate technology problems (although there is that too) and more about being a good influence, convincing the rest of the technical team why Thing A should be solved with plan Foo and not plan Bar. There is a lot of solving for other people motivations, finding the right message for each audience while still working towards the business goals at the end.&lt;/p&gt;

&lt;h3 id=&quot;whats-next&quot;&gt;What’s next&lt;/h3&gt;
&lt;p&gt;The past year and a half working as a principal engineer have been eye opening into all the work that goes into getting a large group of humans all rowing in the same direction for the business to achieve planned goals. It is no small feat that is further complicated with our ever growing, ever complex systems of scale. I do not know whether my experience as a principal engineer will some day translate into becoming an engineering manager someday. Maybe it will. I do know that if I ever made the move to managing people that my time as a principal engineer has taught me plenty that I may have missed out on when focusing only on senior engineer competencies.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Many thanks to &lt;a href=&quot;https://twitter.com/whereistanya&quot;&gt;Tanya Reilly&lt;/a&gt; and &lt;a href=&quot;https://twitter.com/log1kal&quot;&gt;Sean Kilgore&lt;/a&gt; for both being great examples and for their excellent feedback on this. And to Camille Fournier for writing that book.&lt;/em&gt;&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2019/01/28/on-being-a-principal-engineer.html&quot;&gt;On Being A Principal Engineer&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on January 28, 2019.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Oh, what a year it’s been]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2018/12/29/oh-what-a-year-it-s-been.html" />
  <id>https://blog.dbsmasher.com/2018/12/29/oh-what-a-year-it’s-been</id>
  <published>2018-12-29T14:00:00+00:00</published>
  <updated>2018-12-29T14:00:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;This year has been a roller coaster.&lt;/p&gt;

&lt;h3 id=&quot;speaking-at-velocity&quot;&gt;Speaking at Velocity&lt;/h3&gt;
&lt;p&gt;This fall I got to speak about lessons learned scaling databases at the Velocity EU conference in London which was a blast. I got to try new things like high afternoon tea with some amazing women in tech and I basically spent the entire conference chatting with many humans I admire and respect. One great result from that is that I now also joined the program committee for Velocity Santa Clara 2019 and I have been reading a LOT of proposals and learning how to evaluate them from people who are far better at this than me 😄&lt;/p&gt;

&lt;h3 id=&quot;growing-the-team&quot;&gt;Growing the team&lt;/h3&gt;
&lt;p&gt;In January 2018 I was boasting that the DB ops team has ‘doubled’. That was accurate. I was a solo DBA for years and in 2017 we hired our second one. But now I am ending 2018 having hired two more engineers and a third who’s starting in January 2019. People always complain that hiring is hard, especially for positions with a specific emphasis like network or databases and it is true. These days, the focus in the field is increasingly on general practitioner positions. And that makes finding people who want to, or already have, some level of specialized experience in managing databases at scale who are also looking to do more of that in an agile organization is even harder. So early in 2018 my manager and I decided that it was time to grow people in the team. We set forth to hire 2 Jr DBA positions where the focus was far more on potential than existing expertise and knowing that it means a lot more mentoring and teaching for me and Bill, my fellow Sr DBA.&lt;/p&gt;

&lt;p&gt;I don’t want to call a success something that is still ongoing but the team has doubled in headcount again and it’s been lots of fun teaching the new team members how we do things and why we do things the way we do. You can read about all I learned about hiring in &lt;a href=&quot;http://blog.dbsmasher.com/2018/06/09/hiring.html&quot;&gt;this post&lt;/a&gt;from earlier this year.&lt;/p&gt;

&lt;h3 id=&quot;personal-growth&quot;&gt;Personal Growth&lt;/h3&gt;
&lt;p&gt;Growing the team was not just about adding headcount. It was time for me to start focusing even less on day to day database operations and try to help the organization scale with more strategic work. I joined the architecture team and became part of the group that gets to help guide the architectural strategy of things we build. This has turned out to be a good thing for more than one reason.&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;I have talked a lot about including those who care and feed the databases as early as possible in the design phase of building a new thing. Having me in the architecture team means I get to see blueprints by teams very early and provide feedback on things that may be traps in the future.&lt;/li&gt;
  &lt;li&gt;My goal career wise is to become a PE II which means working closer with those who are in that position right now. That is the architecture team.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 id=&quot;whats-to-come&quot;&gt;What’s to come&lt;/h3&gt;
&lt;p&gt;2019 is already going to be a very busy year. Late in 2018 there has been a &lt;a href=&quot;https://www.twilio.com/blog/twilio-to-acquire-sendgrid&quot;&gt;bit of news&lt;/a&gt; about SendGrid which I am sure when the time comes will bring interesting new things to solve.
We are also in the process of a &lt;a href=&quot;https://www.sdxcentral.com/articles/news/sendgrid-re-architects-infrastructure-work-aws/2017/10/&quot;&gt;major rearchitecture&lt;/a&gt; to move our stack to AWS. For databases especially that is both a time of so many choices but also a lot of learning how these new to us data stores actually behave under the SendGrid load.&lt;/p&gt;

&lt;p&gt;This year has been great but it was not without lots of help. People like &lt;a href=&quot;https://mobile.twitter.com/log1kal&quot;&gt;Sean Kilgore&lt;/a&gt;, &lt;a href=&quot;https://mobile.twitter.com/tekbuddha&quot;&gt;John Martin&lt;/a&gt; being mentors and sounding boards for me at work (and for non work things too). Friends like &lt;a href=&quot;https://mobile.twitter.com/nicolefv&quot;&gt;Dr Nicole Forsgren&lt;/a&gt;, &lt;a href=&quot;https://mobile.twitter.com/alicegoldfuss&quot;&gt;Alice Goldfuss&lt;/a&gt;, and &lt;a href=&quot;https://mobile.twitter.com/clynnexx&quot;&gt;Connie-Lynne Villanni&lt;/a&gt; being my back channel and letting me vent about all sorts of things. And many more who have taught me lots this year.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2018/12/29/oh-what-a-year-it-s-been.html&quot;&gt;Oh, what a year it’s been&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on December 29, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[On making connections]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2018/09/26/making-connections.html" />
  <id>https://blog.dbsmasher.com/2018/09/26/making-connections</id>
  <published>2018-09-26T14:00:00+00:00</published>
  <updated>2018-09-26T14:00:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;&lt;em&gt;This post is by request from Baron Schwartz who has &lt;a href=&quot;https://goo.gl/2w3Uak&quot;&gt;written about&lt;/a&gt; forming connections within his company, Vivid  cortex&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I have been at SendGrid for nearly 7 years now and have been working remotely for 4 of those years. While SendGrid is not a ‘remote first’ company, it was distributed in multiple locations from the very beginning. Since the early days, leadership at SendGrid believed that we needed a way to form human connections as a team across the different office locations. Here are how this evolved over time and what things have remained the same.&lt;/p&gt;

&lt;p&gt;In the early years this was an annual event. But we were also a small company focused on growth so we had a very informal travel policy, there was very little resistance to anyone booking a trip to Colorado or California and visiting an office then later expensing that trip. As we grew larger, the inter-office travel budget grew dramatically. At about 4 years ago, with these full company all hands already happening twice a year, it was time to set some guidance around inter office travel. While office visits outside the semi annual all hands can still happen, they need to have a justified business reason in&lt;/p&gt;

&lt;p&gt;It was important for us to make at least one of those all hands away from all the offices. Yes it is more expensive. But I have found that I created the strongest bonds with coworkers not in my local office in those ‘off the grid’ multi day all hands where I get to have dinner with new coworkers, get to exchange inside stories of issues we are facing, projects in progress. I have more than once formulated plans to deal with issues in those evening conversations with coworkers I had not met in person yet until these gatherings. 
Our early year kickoffs have been mostly in Mexico where it is actually pretty well priced getting all inclusive accommodation for everyone with many nearby options for team building outings when the schedules have allowed. Travel time became a concern though as the company size grew and flying so many people became a scaling problem (yes that is a theme when you work at SendGrid, scaling) so now we have made one of the all hands local in the US but not in any of the offices and one of them near our HQ in CO the latter still has plenty of social events in the evening for those coming in from out of town to still socialize.&lt;/p&gt;

&lt;p&gt;The schedule of these events is an area that has changed dramatically over the years. The early day kickoffs used to have only one day of planned meetings and the rest was left unscheduled. The idea there was to encourage socializing, leave room for teams that might be distributed to plan team building events. But as the company grew, the unscheduled time became just lounging by the pool, hanging by the pool bar time and while those are fun too, it is hard to justify for an entire day and a half for a company of a few hundred people. Now our kick offs are more structured. We still have the full first day of presentations from all parts of the organization, recapping the past year and presenting goals for the new one. The second full day is spent in breakout sessions which vary from brainstorming sessions to start the planning of the new year goals, to teaching sessions where principal engineers can do what very much resembles conference talks, showing the newer engineering team members the thought process behind some of our architecture, best practices, how to use certain tools, etc.&lt;/p&gt;

&lt;p&gt;When SendGrid had its first kickoff, the company was 65 people. Total. We are now approaching 500 employees. Like everything we do, these all hands gatherings had to also scale. Let me start by saying that I am always in awe of these events at the skill and hard work by our office managers team , corporate IT and internal event planning teams who start planning these many, many months in advance. Scouting locations, securing necessary internet infrastructure which was crucial when these events used to be in Mexico and as the company grew it became necessary to cater all the meals and that became part of the planning too. I work with an amazing event planning team and I am always in awe of what they accomplish planning these gatherings for an ever growing number of people.&lt;/p&gt;

&lt;p&gt;One of the things that did not change is one rule. No family members come on these trips. We have other social events that are open to all (summer picnics, holiday parties) but this semi annual all hands is about company strategy and team building and those work best without the added work of traveling with one’s family. It’s a difficult choice and I’ve had to miss some of these events due to prior family plans but ultimately i do get the greatest benefit and come back more energized for our plans when I’ve enjoyed the informal chats with coworkers. But like everything, absolutes are not effective. So when people have to miss these events due to any number of reasons, we keep recordings of all the large strategy meetings and post them to the company drive along with slide decks.&lt;/p&gt;

&lt;p&gt;You may think that this post contradicts &lt;a href=&quot;http://blog.dbsmasher.com/teams/2018/04/23/remote.html&quot;&gt;previous writing&lt;/a&gt; by me about how I view distributed teams and the word ‘Remote employee’. I would say this is actually the other side of that story. I have always valued these gatherings when I still worked in The Grid’s SoCal office but I have seen the value in them even more when life required me to move away from a daily commute to the office. I do not take credit for any of planning or execution of these events but it is the investment the company puts in them and constantly improving them with team alignment in mind that has made my journey as an employee who is not in the office everyday not just continue but thrive in more responsibility, more involvement in project planning and more leadership within the company. These connective tissue exercises, with deliberate schedules and thought behind every one of them, are very important for rejuvenating connections for all employees and even more valuable when we do not have the day to day human connection we all crave.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2018/09/26/making-connections.html&quot;&gt;On making connections&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on September 26, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Thoughts on the new State of Devops report]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2018/08/31/accelerate-state-of-devops.html" />
  <id>https://blog.dbsmasher.com/2018/08/31/accelerate-state-of-devops</id>
  <published>2018-08-31T10:46:00+00:00</published>
  <updated>2018-08-31T10:46:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;This week, the latest &lt;a href=&quot;https://cloudplatformonline.com/2018-state-of-devops.html&quot;&gt;State of DevOps&lt;/a&gt; report by the team at &lt;a href=&quot;https://devops-research.com&quot;&gt;DORA&lt;/a&gt; was released. There is a lot of conversations in various social media platforms on what makes an engineering team ‘successful’ and how to increase feature velocity in tech companies and the DORA report does not just rely on anecdotes but surveys the field using rigorous scientific methods and statistical constructs to back or debunk common practices with real data. Because of the strong scientific backbone of their findings, I have been following the report for the last few years. If you have not, the book &lt;a href=&quot;https://www.amazon.com/Accelerate-Software-Performing-Technology-Organizations-ebook/dp/B07B9F83WM/ref=sr_1_1?ie=UTF8&amp;amp;qid=1535649522&amp;amp;sr=8-1&amp;amp;keywords=accelerate&quot;&gt;Accelerate&lt;/a&gt; is a great way to learn the findings of the previous four years.&lt;/p&gt;

&lt;h3 id=&quot;database-management-as-a-part-of-team-performance&quot;&gt;Database management as a part of team performance&lt;/h3&gt;
&lt;p&gt;One of the new things in this year’s report is expanding behaviors by high performing teams to more than just ‘development’ practices. This year expands the view on high performing engineering organizations to how these organizations handle database management.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;We discovered that good communication and comprehensive configuration management that includes the database matter. Teams that do well at continuous delivery store database changes as scripts in version control and manage these changes in the same way as production application changes.
I have &lt;a href=&quot;http://sysadvent.blogspot.com/2016/12/day-2-dbas-priesthood-no-more.html&quot;&gt;written in the past&lt;/a&gt; about breaking DBA team barriers, about the DBA stereotype becoming stale and not really helpful to overall organization goals but I admit my opinions had been till now both anecdotal and strongly biased by my experiences in my own career. In my conversations with other organizations it seemed the more common practice was still to either assign a team (or worse, one person) the task of care and feeding of the databases powering the product and to keep the database changes needed for feature releases as last and completely manual because ‘databases are scary’.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So when the work by DORA was finally released I was delighted to see that this year included data and analysis on the correlation between team performance and how to handle database change management. And it seemed my gut feeling on where we should be going as an industry was actually supported by the data. High performing teams maintained as much as they can about their databases in version control just as they do with feature code. Schema definitions, scripts that do administrative tasks and even code that applies configuration management to the databases were all in version control in high performing and elite teams.&lt;/p&gt;

&lt;p&gt;This practice provides a track record of changes and a way to apply peer review to what is being changed. It also allows for treating the database as part of the larger deploy process. It becomes easier to at least provide visibility to team into when and at what stage is their requested change which allows for less communication overhead when planning deploy timelines. With everything but the data itself in version control, teams can also start applying automation and provide feature teams with autonomy over their data-stores which is also greatly helps team overall velocity.&lt;/p&gt;

&lt;h3 id=&quot;the-role-of-outsourcing&quot;&gt;The role of outsourcing&lt;/h3&gt;
&lt;blockquote&gt;
  &lt;p&gt;Analysis shows that low-performing teams are 3.9 times more likely to use functional outsourcing (overall) than elite performance teams, and 3.2 times more likely to use outsourcing of any of the following functions: application development, IT operations work, or testing and QA. This suggests that outsourcing by function is rarely adopted by elite performers.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The report also makes an assertive point about the effort to outsource development by lower performing teams. It even goes as much as calling them ‘misguided performers ’. This, to me, is directed at efforts to just ‘throw money at’ team velocity without much thought into how to use such an engagement. If you are part of a decision on whether to outsource or not, it is very important from the outset to have some internal agreement on things such as&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;What is the expected benefit&lt;/li&gt;
  &lt;li&gt;What is the definition of success&lt;/li&gt;
  &lt;li&gt;How to measure said success, what agreed upon metrics are we using later to reevaluate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now, I have nothing against consultants. I have worked with some of the best consultant DBAs in the field and have also learned a lot from them. My team has also in the past employed contractors from so called “DevOps for hire” companies. But I have also seen consultant engagements go very very terribly in both cost and morale hits to the in house engineers and the crucial difference has been how we incorporate these consultants into the team just as one would a new full time employee. Hiring consultants to become a new silo away from the rest of the engineering organization to solve problems in a vacuum never, ever, works and somehow executives in many companies still think that is a sane thing to do with the company resources. The Accelerate book goes into more detail into how to use external resource successfully towards team velocity and this year’s report solidifies this with more recent data showing low performing teams more likely to fall in the outsourcing trap as a supposed ‘quick fix’.&lt;/p&gt;

&lt;h3 id=&quot;cloud-infrastructure&quot;&gt;Cloud infrastructure&lt;/h3&gt;
&lt;blockquote&gt;
  &lt;p&gt;How you implement cloud matters
A few years ago, a lot of companies still saw using platforms like AWS or Azure as a way to have more outages and spend a lot more money on operating costs of running their product. Now we seem to have swung the other way since these platforms have become a lot more stable but there still seems to be a presumption that using AWS will magically increase feature release velocity but that it is also a quick thing a company can just swing in a quarter.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is important to internalize some important facts about running infrastructure in cloud platforms&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Companies moving to a cloud platform these days look at it as a way to increase feature velocity. That is clearly stated as a finding from the report data and it is an improvement in the field’s understanding of what advantages running in the cloud provides&lt;/li&gt;
  &lt;li&gt;Now that we established &lt;em&gt;why&lt;/em&gt; you would move your infrastructure, you cannot, &lt;em&gt;cannot&lt;/em&gt; forklift what you already built to a cloud provider and magically expect velocity or even cheaper costs. In fact, it will most definitely cause the opposite of both without deliberate thought into your architecture design. You have to approach a move from on premise to cloud to include some thought into re-architecture of what you are moving. Yes including databases.&lt;/li&gt;
  &lt;li&gt;Use the advantages in the cloud platform you are moving to. Use managed services wherever possible to refocus your engineers’ time to only focus on what makes your product better.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Does this mean that you will no longer need ops people or DBEs? Absolutely not. But the roles of operations engineers and database engineers, like everything else, change over time. This is something that needs to be a separate and longer post but if your team transforms the role of a DBE away from manually running databases and towards helping scale your data architecture, you will find that using managed datastores will actually help free those engineers to provide better value for your business and make up for more than the cost of Aurora or DynamoDB.&lt;/p&gt;

&lt;p&gt;There are so many other gems that make one go “aha!” in the report on what high performing engineering teams actually &lt;em&gt;do&lt;/em&gt; vs what they &lt;em&gt;think&lt;/em&gt; makes them successful. It is pretty great seeing these reports annually being supported by solid science and I very much appreciate the work &lt;a href=&quot;https://twitter.com/nicolefv&quot;&gt;Dr. Nicole Forsgren&lt;/a&gt;, &lt;a href=&quot;https://twitter.com/jezhumble&quot;&gt;Jez Humble&lt;/a&gt; and &lt;a href=&quot;https://twitter.com/RealGeneKim&quot;&gt;Gene Kim&lt;/a&gt; put into it year after year.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2018/08/31/accelerate-state-of-devops.html&quot;&gt;Thoughts on the new State of Devops report&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on August 31, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Hiring]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2018/06/09/hiring.html" />
  <id>https://blog.dbsmasher.com/2018/06/09/hiring</id>
  <published>2018-06-09T10:46:00+00:00</published>
  <updated>2018-06-09T10:46:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;Hiring is hard. It really is. You are looking to add a new human to your team who will hopefully work with you for many months to come but you have about an hour, at best, to develop some strong opinions about how they will perform in the future. We spend way more time than that on vetting vendors, don’t we?
I recently worked with my manager to double the headcount of my team which was both daunting and very illuminating. By sharing how it went and things I have learned along, I hope I can provide some takeaways that can help you when you hire people in the future.&lt;/p&gt;

&lt;h2 id=&quot;but-i-am-not-a-manager&quot;&gt;But I am not a manager!&lt;/h2&gt;
&lt;p&gt;Before we even get into the hiring cycle, you may ask: But I am not a manager! Why am I even getting involved in this? My time is better spent writing code!
Well… the idea that an engineer’s job is solely to churn out code is a myth, and it’s an unhealthy one.
Unless you don’t plan on ever talking to those new teammates, you should be jumping on the chance to help with hiring the people you are about to work with. We spend more time during the week interacting with team members than family members sometimes. Shared core values and complementary problem solving styles can take a team from marginally functional to highly functional and performing strongly. I am lucky to be in an organization that recognizes the importance of senior ICs input when expanding its teams and I was eager to help every step of this process.&lt;/p&gt;

&lt;h2 id=&quot;expectations&quot;&gt;Expectations&lt;/h2&gt;
&lt;p&gt;Before we even published the openings on my team, we needed to get on the same page as to what we expect. This is important because this was the first time either my company or I have hired level 1 DBAs. It is, in some ways, easier to know what the job opening will look like and list its competencies when the need is for a role that presumes previous experience. Level 1, in this case, was closer to associate and needed to focus a lot more on potential while still setting some as yet unknown level of general technical expertise that got us candidates who at least had some idea what a database is.
That turned out harder than I thought. My manager and I started pairing on a Google doc as the basis of the job description and we focused on a few things.
For junior level positions, we avoided listing a bunch of technologies as ‘requirements’. Literally the only things in there were “you have some history using MySQL/a database” and “You have written code before”. The idea here is that this is a role that will already involve a lot of learning. To come into it with a lot of expectation of previous experience will not only be lazy on my part as the future mentor of these employees but would out of the gate exclude any URM candidates who may have genuine interest in becoming DBAs but lack the disposable time/income to self learn these things.
We made sure to explicitly mark ‘nice to have skills’ as ‘bonus skills’. &lt;a href=&quot;https://hbr.org/2014/08/why-women-dont-apply-for-jobs-unless-theyre-100-qualified&quot;&gt;Men apply for a job when they meet only 60% of the qualifications, but women apply only if they meet 100% of them and I was aware of that.&lt;/a&gt; The DBA field, much like the rest of tech, has its own issues around women advancing and lasting in the industry as much as men. I wanted at this entry level position not to compound the problem by inadvertently chasing away potential candidates through making a list of skills that is implicitly sometimes ‘must have’ and sometimes ‘bonus’
Finally, it was important in the job description, upon which our public post was based, to describe more of what the person in this job will do. I was deliberate early on to emphasize the need for someone who will collaborate and work with engineers closely and talk to people. There is a stereotype around DBAs that I did not want to ignore or gloss over and I felt it was important to be upfront early on in the process that this DBA team will not be operating in that manner.
We also made sure to pass the job description through a couple of tools that analyze the language for gender bias. &lt;a href=&quot;http://gender-decoder.katmatfield.com&quot;&gt;One of those&lt;/a&gt; is open source and can be used for free. This was important to make sure we weren’t inadvertently projecting gender stereotypes in how we described the people we were looking for to fill these roles.&lt;/p&gt;

&lt;h2 id=&quot;screening&quot;&gt;Screening&lt;/h2&gt;
&lt;p&gt;We had planned an interview team for these positions. But my hiring manager was the one tasked with screening incoming resumes and doing initial phone screens. This did not have to mean he had to be the sole authority on who comes through the door. In fact, in our conversations about how the screening was going were already shaping and fine tuning what we wanted the roles to be and therefore what kinds of questions to ask in later phases of the interview process.
A lot of things that many people talk about as signals for screening and filtering candidates proved not as useful as we went through resumes of candidates:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Github profiles. That was never even a question. And if you follow me on Twitter you may have caught me &lt;a href=&quot;https://twitter.com/dbsmasher/status/1004933313310572544&quot;&gt;rant&lt;/a&gt; about &lt;a href=&quot;https://twitter.com/dbsmasher/status/1004935870577717249&quot;&gt;this&lt;/a&gt; after someone privileged declared it ‘the new resume’. Let me reiterate here: Whether for a senior role or a more beginner role, personal projects and GitHub activity are a strong indicator of candidate’s free time which is really not an indicator of how well they will collaborate with the larger ops org, delivery engineers or what style of problem solving in our existing infrastructure they will use.&lt;/li&gt;
  &lt;li&gt;Degree or the university they graduated from. Sure, a Bachelor’s degree is a nice to have. But I know too many very good engineers with no college degree to presume it is an indicator of someone’s ability to learn.&lt;/li&gt;
  &lt;li&gt;Tech stack previous experience. This one is a useful signal, to a degree, for senior roles. But when it came to level I hiring where I already expect the hires to need training, it would have been disingenuous to then filter on whether they already knew what I am expected to teach. What matters here is showing hunger to learn.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;interviews&quot;&gt;Interviews&lt;/h2&gt;
&lt;p&gt;The interview team selection also took some thought. I do not believe in DB ops being a silo that is only focused on managing databases. Sure, we still have a good amount of that work at hand as business responsibilities of Data Operations, but I was hiring people who would also learn a lot about this job for the long run and I knew I wanted to stress the collaborative nature of how this team is supposed to operate from day one.
On that note, the on site interview team needed to be more than just me and the hiring manager. We enlisted the team’s project manager to get a sense of candidates’ value alignment and ability to balance unplanned work from outside the team with planned projects. We worked with our People Ops team to prepare relevant score cards in our hiring tool so that when we sit down and have a debrief about a candidate, it is a more productive conversation and we can all come together with a recommendation for the VP of Ops who had the final say before sending an offer.&lt;/p&gt;

&lt;h2 id=&quot;filtering-out-brilliant-jerks&quot;&gt;Filtering out &lt;a href=&quot;http://www.brendangregg.com/blog/2017-11-13/brilliant-jerks.html&quot;&gt;‘brilliant jerks’&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;I know some companies like to do panel interviews. I get the rationale behind it. The most expensive resource is your people’s time and many see it as a way to be efficient. But I am also not a fan of panel interviews. In my experience, and from talking to a number of other women in tech who have been in hiring positions, panels are not very effective at detecting interpersonal issues with candidates that can cause problems down the line. Let me explain how that plays out.
Sexism, racism and bias in tech are not a secret. If you are reading this and still believe that hiring in tech is a meritocracy based solely on skills and never about impressions, implicit and explicit biases then you should go read some research &lt;a href=&quot;http://geekfeminism.wikia.com/wiki/Geek_Feminism_Wiki&quot;&gt;here&lt;/a&gt; and &lt;a href=&quot;http://projectinclude.org&quot;&gt;here&lt;/a&gt; then come back to the rest of this post.
Now that we have set this baseline, it is a natural deduction that when you are tasked by helping your team grow, it is important to make value alignment an important part of the process. A 1:1 setting for interviews is far more effective in showing how candidates will behave around team members who are from under represented groups (URM) in tech. Sure, there will always be the more egregious examples of candidates so full of their ego that they talk down to a woman on the team in front of everyone else but that will only happen if you are lucky. What you should strive to detect for is the more sinister attitude of potential team members who could hurt the team morale using a thousand tiny cuts to their URM team members. The cost of a toxic member coming on your team is far greater to your team and the organization than the cost of some more interviews to find candidates whose values align with yours.
This is not to say that the on site interview process was solely focused on interpersonal skills, we also needed to gauge for problem solving skills from the technical standpoint. Database engineering is still a job that has on call requirements and it is still very possible to get asked to debug issues in production before anyone is even sure if it is a database bottleneck or something else entirely so we needed to screen the candidates for their ability to inspect systems and to show how they approach debugging on live systems which can often feel like a murder mystery. Some teams like to prepare test code or a service they write specifically for these exercises but in my case it was as simple as a verbal hypothetical and a number of follow-ups as the candidate tells me what they would look at and what they expect to find. For me, what I watch out for here is persistence, curiosity and demeanor under pressure far more than whether they guess what the problem actually is in this hypothetical. When an outage situation escalates to my team it tends to mean it is already an all hands on deck situation, so it is very important to see that the people on my team will be calm under such pressure.&lt;/p&gt;

&lt;h2 id=&quot;debriefs&quot;&gt;Debriefs&lt;/h2&gt;
&lt;p&gt;After each onsite interview, the entire interview team along with the people ops recruiter responsible for this hire huddle together to discuss our score cards. This is where discrepancies in the candidate behavior we may have encountered are brought up and discussed. This candidate was respectful and charming with the hiring manager but kept interrupting and talking down to the project manager? Nope. The next candidate showed frustration at scenarios where we discussed feature planning with product and delivery engineers? Also no. That senior DBA candidate made an off hand comment about how he never wants developers to touch his database? Also no. Usually in these debriefs we look for strong yes answers on candidates or we discuss the interviews we did with that person some more. We talk more about reasons we may not hire a person because it is important to surface these concerns early. Many times I have seen teammates come into these debriefs as a weak yes on a candidate only to become a solid no after they go over notes from the other people on the interviewing team and that is why we do these debriefs.&lt;/p&gt;

&lt;h2 id=&quot;things-i-learned&quot;&gt;Things I learned&lt;/h2&gt;
&lt;p&gt;My road to where I am now in my career was full of serendipitous encounters, a chain of things that if they hadn’t happened, I wouldn’t be where I am today; If I hadn’t found a job as the DBA’s intern in college, if my first job in tech wasn’t with such a small team that they’d let me, 6 months out of college, debug issues with the billing code and the database.
This process of hiring new level 1 DBAs, knowing that I was about to have a lasting impact on two other people’s careers was a humbling thing and it made me take note of a lot of the advantages I had in my journey. It brought back memories to spending days in MySQL documentation learning about all sorts of configuration and commands, then later spending days in Chef documentation trying to understand how the hell attributes work.
During the initial screenings and after a couple of initial on site interviews, my manager and I had some honest conversations about what preexisting knowledge we were truly testing for and whether that was fair. Yes! That was already discussed when writing the job description, but it was important to keep revisiting that because it is easy to want the knowledge to be already there. As he, myself and my fellow senior DBA on the team talked about it more, we determined that we were looking far more for potential than expertise and that quickly made a large chunk of prepared interview questions carry a lot less weight in how we scored candidates. I started using the ‘technical phone screen’ as a way to motivate the candidates’ curiosity and hunger. The two engineers we hired showed hunger for reading material in the initial screening with me and when they came to the on site with prepared questions based on things they already read and learned, I knew that they had the curiosity to learn about this stuff to make it past the initial steep turning curve of the database toolset.&lt;/p&gt;

&lt;h2 id=&quot;a-team-to-last&quot;&gt;A team to last&lt;/h2&gt;
&lt;p&gt;My team has grown from a team of one (me) to now four people strong. This has been great not just for sheer team velocity and being able to service a large engineering organization but has also given me a ton of room to grow in my career. I have benefitted a lot from showing my fellow senior DBA how I have been doing things for a while, he has taught me a lot and now we have 2 level 1 DBAs who are giving us the opportunity to mentor them and they are excited to do all the tasks we were ready to grow out of. The biggest lesson to me in this past 18 months of growing the team has been than the most important thing is hiring people who have empathy, humility to learn, and curiosity. A good friend and my former manager recently sent me a link this post and this line stood out to us in the email discussion.&lt;/p&gt;
&lt;blockquote&gt;
  &lt;p&gt;The most productive teams I’ve been on are the ones who are pulling each other over these walls in the battle against real-world problems.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Managing infrastructure, data stores or otherwise, is a team sport. And that is why the process for joining the team has to optimize for team collaboration and psychological safety.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2018/06/09/hiring.html&quot;&gt;Hiring&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on June 09, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[How to choose a data store for the new shiny thing]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2018/05/12/how-to-choose-a-data-store-for-the-new-shiny-thing.html" />
  <id>https://blog.dbsmasher.com/2018/05/12/how-to-choose-a-data-store-for-the-new-shiny-thing</id>
  <published>2018-05-12T04:43:00+00:00</published>
  <updated>2018-05-12T04:43:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;&lt;em&gt;Note: This is post first appeared as part of the SysAdvent series of December of 2017&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Databases can be hard. You know what’s harder? Choosing one in the first place. This is challenging whether you are in a new company that is still finding its product/market fit or in a company that has found its audience and is simply expanding the product offering. When building a new thing, one of the very first parts of that design process is what data stores should we use and should that be a single or a plural? Should we use relational stores or do we need to pick a key value store? What about time options? Should we also sprinkle in some distributed log replay? So. Many. Options…&lt;/p&gt;

&lt;p&gt;I will try in this article to describe a process that will hopefully guide that decision and, where applicable, explain how the size and maturity of your organization can impact this decision.&lt;/p&gt;

&lt;h2 id=&quot;baseline-requirements&quot;&gt;Baseline requirements&lt;/h2&gt;

&lt;p&gt;Data is the lifeblood of any product. Even if we’re planning in the design to use more bleeding edge technology to store the application state (because MySQL or Postgres aren’t “cool” anymore), whatever we choose is still a data store and hence requires that we apply rigor when making our selection. The important thing to remember is that nothing is for free. All data stores come with compromises and if you are not being explicit about what compromises you are taking as a business risk, you will be taking unknown risk that will show itself at the worst possible time.&lt;/p&gt;

&lt;p&gt;Your product manager is unlikely to know or even need to care what you use for your data store but they will drive the needs that shrink the option list. Sometimes even that needs some nudging by the development team, though. Here is a list of things you need to ask the product team to help drive your options:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Growth rate - How is the data itself or the access to it expected to change over time?&lt;/li&gt;
  &lt;li&gt;How will the billing team use this new data?&lt;/li&gt;
  &lt;li&gt;How will the ETL team use this data?&lt;/li&gt;
  &lt;li&gt;What accuracy/consistency requirements are expected for this new feature?&lt;/li&gt;
  &lt;li&gt;What time span for that consistency is acceptable? is post processing correction acceptable?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;find-the-context-that-is-not-said&quot;&gt;Find the context that is not said&lt;/h2&gt;

&lt;p&gt;The choice of the data store is not a choice reserved for the DBA or the Ops team or even just the engineer writing the code. For an already mature organization with a known addressable market, the requirements that feed this decision need to come from across the organization. If the requirements from the product team fit a dozen data stores, how do you determine requirements not explicitly called out? You need to surface unspoken requirements as soon as possible because it is the road to failed expectations down the line. A lot of implied things can make you fail in this ‘too many choices’ trap. This includes but is not limited to:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Incomplete feature lists&lt;/li&gt;
  &lt;li&gt;Performance requirements that are not explicitly listed&lt;/li&gt;
  &lt;li&gt;Consistency needs that are assumed&lt;/li&gt;
  &lt;li&gt;Growth rate that is not specified&lt;/li&gt;
  &lt;li&gt;Billing or ETL query needs that aren’t yet available/known&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are all possible ways that can leave an engineering team spinning their wheels too long vetting a long list of data store choices simply because the explicit criteria they are working with are too permissive or incomplete.&lt;/p&gt;

&lt;p&gt;For more ‘greenfield’ products, as i mentioned before, your goal is flexibility. So a more general purpose, known quality, data store will help you to get closer to a deliverable, with the knowledge that down the line, you may need to move to a datastore that is more amenable to your new scale.&lt;/p&gt;

&lt;h2 id=&quot;make-your-list&quot;&gt;Make your list&lt;/h2&gt;

&lt;p&gt;It is time to filter potential solutions by your list of requirements. The resulting list needs to be not more than a handful of possible data stores. If the list of potential databases you can use is more than that then your requirements are too permissive and you need to go back and find out more information.&lt;/p&gt;

&lt;p&gt;For younger, less mature companies, data store requirements is the area of the most unknowns. You are possibly building a new thing that no one offers just yet and so things like total addressable market size and growth rate may be relatively unknown and hard to quantify. In this case what you need is to not constrain yourself too early in the lifetime of your new company by using a one trick pony datastore. Yes, at some point your data will grow in new and unexpected ways but what you need right now is flexibility as you try to find your market niche and learn what the growth of your data will look like and what specific scalability features will become crucial to your growth.&lt;/p&gt;

&lt;p&gt;If you are a larger company with a growing number of paying customers, your task here is to shrink the option list to preferably data stores you already have and maintain. When you already have a lot of paying customers, the risk of adding new data stores that your team is not familiar with becomes higher and, depending on the context of the data, simply unacceptable. Another thing to keep in mind is what tooling already exists for data stores and what would adopting a new one mean as far as up front work your team has to do. Configuration management, backup scripts, data recovery scripts, new monitoring checks, new dashboards to build and get familiar with. The list of operational cost of a new data store, regardless of risk, is not trivial.&lt;/p&gt;
&lt;h2 id=&quot;choose-your-poison&quot;&gt;Choose your poison&lt;/h2&gt;

&lt;p&gt;So here is a badly kept secret that DBAs hold on to. Databases are all terrible at something. There is even a whole theorem about that. Not just databases in the traditional sense, but any tech that stores state will be horrible in a way unique to how you use it. That is just a fact of life that you better internalize now. No, I am not saying you should avoid using any of these technologies, I am saying keep your expectations sane and know that YOU and only you and your team ultimately own delivering on the promises you make.&lt;/p&gt;

&lt;p&gt;What does this mean in non abstract terms? Once you have a solid idea what data stores are going to be part of what you are building, you should start by knowing the weaknesses of these data stores. These weaknesses include but are not limited to:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Does this datastore work well under scan queries?&lt;/li&gt;
  &lt;li&gt;Does this datastore rely on a gossip protocol for data replication? if so, how does it handle network partitions? How much data is involved in that gossip?&lt;/li&gt;
  &lt;li&gt;Does this datastore have a single point of failure?&lt;/li&gt;
  &lt;li&gt;How mature are the drivers in the community to talk to it or do you need to roll your own?&lt;/li&gt;
  &lt;li&gt;This list can be &lt;em&gt;huge&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thinking through the weaknesses of the potential solutions still on your list should knock more options off the list. This is now reality meeting the lofty promises of tech.&lt;/p&gt;

&lt;h2 id=&quot;spreadsheet-and-bake-off&quot;&gt;Spreadsheet and Bake off!&lt;/h2&gt;

&lt;p&gt;Once your list of choices is down to a small handful, it is time to put them all in a spreadsheet and start digging a little deeper. You need a pros column and a cons column and at this point, you will need to spend some time in each database documentation to find out nitty gritty details on how to do certain tasks. If this is data you expect to have a large growth rate, you need to know which of these options is easier to scale out. If this is a feature that does a lot of fuzzy search, you need to know which datastore can handle scans or searching through a large number of rows better and with what design. The target at this stage is to whittle down the list to ideally 2 or 3 options via documentation alone because if this new feature is critical enough to the company success, you will have to benchmark all three.&lt;/p&gt;

&lt;p&gt;Why benchmark you say? Because no 2 companies use the same datastore the same way. Because sometimes documentation implies caveats that only gets exposed in other people’s war stories. Because no one owns the stability, the reliability and the predictability of this datastore but you.&lt;/p&gt;

&lt;p&gt;Design your benchmark in advance. Ideally, you set up a full instance of the datastores in your list with production level specifications and produce test data that is not too small to make load testing useless. Make sure to not only benchmark for ‘normal load’ but also to test out some failure scenarios. The hope is that through the benchmark, you can find out any caveats that are severe enough to cause you to revisit the option list now instead of later when all the code is written and you are now at the fire drill phase with a lot of time and effort committed to the choice you made.&lt;/p&gt;

&lt;h2 id=&quot;document-your-choice&quot;&gt;Document your choice&lt;/h2&gt;

&lt;p&gt;No matter what you do, you must document and broadcast internally the method by which you reached your choice and the alternatives that were investigated on the route to that decision. Presuming there is an overarching architecture blueprint of how this new feature will be created and all its components, you make sure to create a section dedicated to the datastore powering this new feature with links to all the benchmarks done to reach the decision the team came to. This is not just for the benefit of future new hires but also for your team’s benefit in the present. A document that people can asynchronously read and develop opinions on provides a way to keep the decision process transparent, grow a sense of best intent among the team members and can bring in criticism from perspectives you didn’t foresee.&lt;/p&gt;

&lt;h2 id=&quot;wrap-up&quot;&gt;Wrap up&lt;/h2&gt;

&lt;p&gt;These steps are not only going to lead to data-informed decisions when growing the business offering, but will also lead to a robust infrastructure and a more disciplined approach to when and where you use an ever growing field of technologies to provide value to your paying customers.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2018/05/12/how-to-choose-a-data-store-for-the-new-shiny-thing.html&quot;&gt;How to choose a data store for the new shiny thing&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on May 12, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[ Encrypting All Our Backups: On Making It To That Finish Line]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2018/05/08/encrypting-all-our-backups-on-making-it-to-that-finish-line.html" />
  <id>https://blog.dbsmasher.com/2018/05/08/encrypting-all-our-backups-on-making-it-to-that-finish-line</id>
  <published>2018-05-08T03:21:00+00:00</published>
  <updated>2018-05-08T03:21:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;&lt;em&gt;Note: This is a repost of a blog post I wrote for Sendgrid’s blog in September of 2017&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A year ago, Sendgrid was working hard towards SOC2 certification. Everyone was involved. There were stories on nearly every delivery team board with a SOC2 tag as we were all looking to be certified by the end of the third quarter. As you can imagine, being the person in charge of databases, there was definitely some work to do for that part of the stack.&lt;/p&gt;

&lt;p&gt;On my task list for this business-wide endeavor was making sure that our backups were encrypted. Since my area of familiarity is DBA tools and knowing that Percona’s xtrabackup already has support for encryption, it was predictable that I would go to that as the first attempt at this task.&lt;/p&gt;

&lt;p&gt;A few important things were in my sights in testing this approach:
Obviously, the backup needed to be encrypted
The overhead to creating the backup needed to be known and acceptable
The overhead to decrypting the backup at recovery time needed to be known and acceptable&lt;/p&gt;

&lt;p&gt;That meant I needed first to be able to track how long my backups take.&lt;/p&gt;

&lt;p&gt;Sendgrid uses graphite for its infrastructure metrics and while the vast majority are sent via sensu, graphite is easy enough to send metrics directly via bash lines– very convenient since the backup scripts are in bash. Note that sending metrics at graphite directly is not super scalable but since these backups run at most once an hour, that was fine for my needs.&lt;/p&gt;

&lt;p&gt;That part turned out to be relatively easy.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;START=$(date +%s)
# do a bunch of backup related things
END=$(date  +%s)
runtime=$(($END-$START))
echo &quot;databases.backups.$db_prefix.full  $runtime `date +%s`&quot; | nc -w 1 $graphite_url 2003
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;To explain what happened in that last line, I send graphite the path of the metric I am sending (make sure that is unique), the metric value, then the current time in epoch format. Netcat is what I decided to use for simplicity, and I give it a timeout of 1 second because, otherwise, it will never exit. the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;graphite url&lt;/code&gt; is our DNS endpoint for graphite in the data center.&lt;/p&gt;

&lt;p&gt;Now that I had a baseline to compare to, we were ready to start testing encryption methods.&lt;/p&gt;

&lt;p&gt;Following &lt;a href=&quot;https://www.percona.com/doc/percona-xtrabackup/LATEST/innobackupex/encrypted_backups_innobackupex.html&quot;&gt;the detailed documentation&lt;/a&gt; from Percona on how to do this I started out by making a key. If you read that documentation page carefully, you may realize something.&lt;/p&gt;

&lt;p&gt;This key is to be passed to the backup tool directly, and it is the same key that can decrypt the snapshot. That is called symmetric encryption and it is, by nature of that same key in both direction, less secure than asymmetric encryption. I decided to continue testing to see if simplicity still makes this a viable approach.
Tests with very small DBs, a few hundred MBs, were successful. The tool works as expected and documented but that was more of a functional test and the real question was “what is the size of the penalty of encryption on our larger DBs?” The more legacy instances at SendGrid had grown to sizes from 1-2 TB to a single 18 TB beast. What I was going to use for the small instances had to also be acceptably operational on the larger ones.&lt;/p&gt;

&lt;p&gt;This is where testing and benchmarks got interesting.&lt;/p&gt;

&lt;p&gt;My first test subject of a considerable size is a database we have that is 1 TB on disk. Very quickly I encountered an unexpected issue. With minimal encryption settings (1 thread, default chunk sizes) I saw the backups fail with this error:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;xtrabackup: error: log block numbers mismatch:
xtrabackup: error: expected log block no. 243782669, but got no. 245879813 from the log file.
xtrabackup: error: it looks like InnoDB log has wrapped around before xtrabackup could process all records due to either log c
opying being too slow, or  log files being too small.
xtrabackup: Error: xtrabackup_copy_logfile() failed.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;At the time these databases used 512MB as the transaction log file size, and this is a fairly busy cluster so those files were rotating almost every minute. Normally this would be noticeable in the DB performance but it was mostly masked by the wonder of solid state drives. Seems like not setting any parallel encryption threads (read: use one) means we spend so much time encrypting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.ibd&lt;/code&gt; files that the innodb redo log rolls from under us making the backup break.&lt;/p&gt;

&lt;p&gt;Let’s try this again with a number of encryption threads. As a first attempt, I tried with 50 threads. The trick here is to find the sweet spot of fast encryption without competing over CPU. I also increased the size of the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_logfiles&lt;/code&gt; to 1 GB each.&lt;/p&gt;

&lt;p&gt;This was a more successful test that I was happy to let brew overnight. For the first few nights, things seemed good. It was time to make a backup that doesn’t grow too much, but box load average during the backup process was definitely showing the added steps.&lt;/p&gt;

&lt;p&gt;However, when I moved on to testing restores, I found that the restore process of the same backups, after adding encryption, had increased from 60 to 280 minutes. That meant a severe penalty to our promised recovery time in case of disaster and we needed to bring that back to a more reasonable timeframe.&lt;/p&gt;

&lt;p&gt;This is where teamwork and simpler solutions to problems shined. One of our InfoSec team members decided to see if this solution can be simplified. So he did some more testing and came back with something simpler and more secure. I had not yet learned about &lt;a href=&quot;https://linux.die.net/man/1/gpg2&quot;&gt;GPG2&lt;/a&gt; and so this became a learning exercise for me as well.&lt;/p&gt;

&lt;p&gt;The good thing about gpg2 is that it supports asymmetric encryption. The way that works is that we create a key pair where there is a private and public parts. The public part is used to encrypt any stream or file you decide to feed gpg2 and the private secret can be used to decrypt.&lt;/p&gt;

&lt;p&gt;The change to our backup scripts to add encryption distilled to this. Some arguments are removed to make this easier to read:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;innobackupex  $backup_path 2&amp;gt;&amp;gt;$log_file | pigz | /usr/local/bin/gpg2 --no-default-keyring --keyring /var/lib/mysql/backup_keyring.gpg --trust-model always --encrypt --digest-algo sha512 --cipher-algo aes256 --compress-level 0 --recipient REDACTED --recipient REDACTED | ssh $remote_user@$remote_server &quot;cat &amp;gt; $remote_path/$date_path/FS_$backup_file&quot; 2&amp;gt;&amp;gt;$local_err

&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;On the other end, when restoring a backup, we simply have to make sure a secret key that is acceptable is in the host’s keyring and then use this command:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;/usr/bin/gpg2 --decrypt $backupfile.gpg &amp;gt; backup_file.xbs
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Since I was new to gpg2 as well, this became a learning opportunity for me. Colin, our awesome InfoSec team member, continued to test backups and restores using gpg2 until he confirmed that the using gpg2 had multiple advantages to using xtrabackup’s built in compression&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;It was asymmetric&lt;/li&gt;
  &lt;li&gt;Its secret management, for decryption, is relatively easily rotated (More on this below)&lt;/li&gt;
  &lt;li&gt;It is tool agnostic, which means any other kind of backup that isn’t using xtrabackup could use the same method&lt;/li&gt;
  &lt;li&gt;It was using multiple cores on decryption which gave us better restore times&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One place where I still see room for improvement is how we handle the secret that can decrypt these files. At the moment we have them in our enterprise password management solution, but getting it from there then using it to test backups is a manual process. Next in our plan is to implement &lt;a href=&quot;https://www.vaultproject.io&quot;&gt;Vault by Hashicorp&lt;/a&gt; and use that to seamlessly, on the designated hosts, to pull the secret key for decryption, then remove it from the local ring so it is easily available for automated tests and still protected.&lt;/p&gt;

&lt;p&gt;Ultimately, we got all of our database backups to comply with our SOC2 needs in time without sacrificing backup performance or disaster recovery SLA. I had lots of fun working on this assignment with our InfoSec team and came out of it learning new tools. Plus it is always nice when the simplest solution ends up being the most suitable for one’s task.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2018/05/08/encrypting-all-our-backups-on-making-it-to-that-finish-line.html&quot;&gt; Encrypting All Our Backups: On Making It To That Finish Line&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on May 08, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA['Remote']]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/teams/2018/04/23/remote.html" />
  <id>https://blog.dbsmasher.com/teams/2018/04/23/remote</id>
  <published>2018-04-23T08:39:00+00:00</published>
  <updated>2018-04-23T08:39:00+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;I never liked that term ‘remote’. Whether in relation to me or to colleagues. It is exclusionary. And betrays a sense of ‘otherness’ that I feel is unhealthy for team cohesion. I understand that not everyone finds it easy or productive to work mostly apart from daily, in person, social interaction. But success is an arm in arm effort and the strategies to build successful teams work regardless of geography if we reward the right behaviors.&lt;/p&gt;

&lt;p&gt;Every now and then, discussions around distributed teams, the hiring practices of a distributed team and all the details that go along with that come up whether in Twitter or in person with friends and co-workers.&lt;/p&gt;

&lt;p&gt;If you follow me on twitter, it won’t be news that I work ‘remotely’ and have been for a few years since a couple of family moves. It has been a mostly successful experience that still goes on. I am not only still on the team, but I have since seen a promotion, have grown the team and am now growing it even more by hiring some level I DBAs to mentor and help grow.&lt;/p&gt;

&lt;p&gt;So yes. It’s been and continues to be a success story. Even in a company that isn’t remote first.&lt;/p&gt;

&lt;p&gt;What I do not talk about often is the story of when I didn’t succeed as a remote employee, resulting in being laid off years ago. And I think it’s important to talk about that as well. Because as gratifying as it may be to claim full credit at my current success, I’d be dishonest to pretend that success was all me.&lt;/p&gt;

&lt;p&gt;The past few years were not the first time I’ve worked from home. However, when I’ve done it before it ended in lack of engagement and ultimately in my value being seen as dispensable, replaceable. I simply didn’t have a team around me that valued communication or the people management apparatus that understood what i did day to day besides just showing my face in meetings.&lt;/p&gt;

&lt;p&gt;So what is it that makes the same person excel or fail miserably at the same thing? Let me share some of my experience around this and how I think it may matter to your organization.&lt;/p&gt;

&lt;h3 id=&quot;the-team&quot;&gt;The team&lt;/h3&gt;
&lt;p&gt;There is never a ‘successful remote employee’. There is always a ‘team that communicates well’ and this goes beyond locality and who’s physically where. I’ve seen Engineering teams in the same office fail spectacularly at basic communication around work in progress or around setting expectations for delivering feature work while other teams with three different timezones producing results efficiently. This is all a roundabout way for me to say that my success is really my team’s success.
There are a number of habits that make my team so good at this and they are worth calling out&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Most conversations happen in chat. Keeps everyone in the loop&lt;/li&gt;
  &lt;li&gt;Joining stand-up from laptops even for those in the office. Everyone is on the same communication channels&lt;/li&gt;
  &lt;li&gt;In meetings where a large portion of the group is in a big room, someone is designated as representative of those calling in. That way we can still get the group’s attention and ask questions if we want to&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are not just behaviors that are useful when not the whole team is local. Life happens, co-workers who are parents have to work from home sometimes, errands, all sorts of things. Allowing for your people to still be part of important discussions while living life is a real sign of wanting life balance for your people and not just talking about it.&lt;/p&gt;

&lt;h3 id=&quot;the-effects-of-human-laziness&quot;&gt;The effects of human laziness&lt;/h3&gt;
&lt;p&gt;Humans are by nature lazy, we do not seek what we do not see. And we make presumptions about things we do not know for a fact. These natural shortcuts our brains take will always show themselves in the systems we design. I have seen it multiple times where teams have developed and deployed software that did not actually meet all the needs of the business only for retrospectives later to declare the ‘root cause’ a failure in communication.&lt;/p&gt;

&lt;p&gt;Noting the skepticism I have towards any mention of ‘root cause’, it amazes me how often organizations will over and over again prove Conway’s law in this specific dynamic. Teams that only communicate to humans in front of them will also fail to involve all stakeholders in decisions in design, will fail to notify other teams when they are about to ratchet up customer involvement in their new beta and will  create products that ultimately are a reflection of how the team itself handles its communication. It has been my experience that teams that put the effort into being good at communication, at writing things down, at making everything explicit and as little as possible implicit, that also find the most alignment and find all their members rowing in the same direction. Having people work from wherever they live at this point becomes the cherry on top.&lt;/p&gt;

&lt;h3 id=&quot;diversity-and-inclusion&quot;&gt;Diversity and inclusion&lt;/h3&gt;
&lt;p&gt;Here’s a harsh truth: If you require hiring in a specific geographic location but pinky swear that you value diversity and inclusion, then you are at best misguided as to how to create diverse teams and how to keep underrepresented groups included.
Intent doesn’t matter when hiring practices have the opposite impact. Let’s say you are a company in the bay. Building a new product that has to solve large scale problems ,you know you need experienced distributed systems engineers. What average number of years experience do you think you will find in the bay?&lt;/p&gt;

&lt;p&gt;If you decide to open the experienced roles to ‘anywhere US’ but none of your Engineering  managers have managed people not local to them before or worse are not interested in that,  you are hiring experienced engineers to see them fail. Remember, People quit managers not companies. Engineering managers who have not in the past honed their emotional intelligence to manage people not necessarily physically near them will not retain experienced engineers who have built the experience to know all the skeletons in your infrastructure. One thing I have learned in the last few years is that backchannels are a strong force in this industry. And people will tell each other whether that ‘Anywhere US’ position company Foo just posted is really something that is setup for success in a team that uses mature communication skills or if it is a Hail Mary because that position was open for nine months in the Bay and no one is applying and now that company is desperate to fill it.&lt;/p&gt;

&lt;p&gt;We are in an industry that has far more demand for talent than is available to hire. And that’s an advantage for us as workers. Somehow I see companies clearly state hiring as a potential challenge and something that’s a competitive edge, yet they hamstring those same efforts out of the gate by requiring that people relocate. Not only does this make hiring harder but it is also not inclusive.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/teams/2018/04/23/remote.html&quot;&gt;&apos;Remote&apos;&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on April 23, 2018.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[2017 in review]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2017/12/25/2017-in-review.html" />
  <id>https://blog.dbsmasher.com/2017/12/25/2017-in-review</id>
  <published>2017-12-25T20:11:42+00:00</published>
  <updated>2017-12-25T20:11:42+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;In some ways it feels like this year flew by, in others like it was as long as an eternity. Looking back, it definitely surprises even &lt;em&gt;me&lt;/em&gt; how many significant events have happened in these 12 months. I know one can say that about pretty much every year but it feels more pronounced this time. Without much ado, here is a recap of how this year has been.&lt;/p&gt;

&lt;h3 id=&quot;team-of-one-no-more&quot;&gt;Team of one no more&lt;/h3&gt;

&lt;p&gt;By the beginning of 2017, we had secured an accepted offer from Bill who was to become Sendgrid’s second in house DBA. While the larger tech ops org had always had my back during maternity leaves and vacations and we’ve always had a strong support team in the folks in Pythian, it was clear that we were overdue for the in house DB Ops team to no longer remain a one woman team.&lt;/p&gt;

&lt;p&gt;This was not just a great thing for the team and Sendgrid, but also tremendously useful for my career growth. I had been running on fumes for a while and started to chafe at my day to day tasks, itching to get more involved in large strategy planning rather than drowning in daily tactical tasks. But to do all that I needed to also learn to let go of some things I have had full control over for while. Now, almost a year later, I have realized that having a trusted partner on my team has led to serendipitous changes like letting engineering teams write the chef cookbooks for managing their new database clusters. This process of ‘letting go’ was a crucial first step to the next big career event in 2017 for me. I was told recently that in the past year I have become more pragmatic. And I think this was the single most unexpected but welcome result of not working in isolation anymore.&lt;/p&gt;

&lt;h3 id=&quot;new-titlea-whole-new-role&quot;&gt;New title…a whole new role&lt;/h3&gt;

&lt;p&gt;It was clear to me in 2016 that i was hitting a ceiling in career growth if I remained a single person team doing tactical work all the time. Once the team grew by 100% (😃), I was able to focus more on higher level planning. Getting more involved in architecture blueprints, writing more blog posts, planning my first ever conference talk (more on all that later).&lt;/p&gt;

&lt;p&gt;By late 2017, I earned my promotion to Principal DBA which, while the next step in the IC career track at Sendgrid, is also a role that is far more leadership than ‘heads down on code’. Since this promotion, my role has now shifted from working on code to lots and lots of planning and reading or architecture blueprints. Yes, it means more meetings..but it also means having a broader impact on how we build or rebuild pieces of our architecture and supporting entire teams. If being a ‘senior engineer’ requires one to be a force multiplier within their team, it feels that principal engineer demands that ten times as much and with a much larger impact across the organisation.&lt;/p&gt;

&lt;p&gt;I am still working on accepting changes that have come with this role change, my calendar is certainly a testament and my standup updates are almost always a mix of ‘Talked to Person X about project Foo and Person Y about project Bar’. Getting my head as an engineer out of classifying these conversations as ‘non work’ and instead recognising them as critical collaboration as part of a distributed engineering organisation is now part of my job. As my dear friend Sean Kilgore said in a  tweet:  “My specialty is random conversations. And all of the gdocs suggestions.” He meant it in jest but I think this is not a bad goal to have in 2018 😃&lt;/p&gt;

&lt;h3 id=&quot;blog-posts&quot;&gt;Blog posts&lt;/h3&gt;

&lt;p&gt;While technical posts such as &lt;a href=&quot;https://sendgrid.com/blog/encrypting-our-backups-making-it-to-that-finish-line/&quot;&gt;How we encrypted our backups&lt;/a&gt; are always fun and rewarding once a project is complete, it seems like the most popular post I wrote this past year was a more personal one on &lt;a href=&quot;https://dbsmasher.com/2017/09/30/on-leadership-vs-management/&quot;&gt;management vs leaderhsip&lt;/a&gt;. I will not rehash here my thoughts on the subject but I will note that, for all the focus many people in tech put on the code and the tools. It is definitely of note that blog posts on the human interactions side of the job that seem to get a lot more attention and spark more discussion. It is a good thing. It is about time we stopped pretending this field doesn’t have human vs human responsibilities.&lt;/p&gt;

&lt;h3 id=&quot;second-tech-conference-first-talk&quot;&gt;Second tech conference, first talk&lt;/h3&gt;

&lt;p&gt;One of the highlights of 2017 was being able to attend and speak at LISA. LISA is one of the oldest tech conferences around and through its selections of chairs and talk chairs, has done a great job at being inclusive to attendees and first time speakers, myself included.&lt;/p&gt;

&lt;p&gt;My talk was about &lt;a href=&quot;https://www.youtube.com/watch?v=Ym408YX2zTA&quot;&gt;working with DBAs in a Devops world&lt;/a&gt; which was in part what I talked about, but also the title was more of a hook to talk about “how to architect products by talking to people”. I wrote a &lt;a href=&quot;https://opensource.com/article/17/10/working-dbas-devops-world&quot;&gt;preview blog post&lt;/a&gt; which was very helpful in shaping exactly what I was going to cover besides the outline of the proposal. I also got lots of help and support from wonderful people like &lt;a href=&quot;http://blog.alicegoldfuss.com&quot;&gt;Alice Goldfuss&lt;/a&gt; and &lt;a href=&quot;https://twitter.com/clynnexx&quot;&gt;Connie-Lynne Villani&lt;/a&gt; in both enouraging me to submit to LISA and at the conference as a a ball of nervous energy until I was done presenting.&lt;/p&gt;

&lt;p&gt;LISA was also my second ever tech conference to attend and I realized that i now much prefer conferences with really diverse attendance. I got to meet so many fellow women in tech and had plenty of awesome conversations that I will cherish for a long time and it was all because of a planning committee that put real work in making sure the conference had a friendly and inclusive environment.&lt;/p&gt;

&lt;h3 id=&quot;ipo&quot;&gt;IPO!&lt;/h3&gt;

&lt;p&gt;January 2018 will mark six years working for Sendgrid. When I joined I had no idea that the business of delivering emails was….a business…and possibly a lucrative one. It was a highlight of this year and possibly the rest of my career being in NYC and, in person, watching the company I have poured years of hard work in go public and have its day in the sun. While this is only a milestone and not a destination, it is definitely a milestone many companies strive to accomplish and it felt good being a part of it.&lt;/p&gt;

&lt;h3 id=&quot;wrap-up&quot;&gt;Wrap-up&lt;/h3&gt;

&lt;p&gt;It has been a &lt;em&gt;very&lt;/em&gt; busy year indeed. I have not yet thought about any personal goals in 2018 but I sure have plenty of them professionally. If it turns out to be half as awesome as 2017 has been, it won’t be too bad. 😉&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2017/12/25/2017-in-review.html&quot;&gt;2017 in review&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on December 25, 2017.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[On leadership vs management]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2017/09/30/on-leadership-vs-management.html" />
  <id>https://blog.dbsmasher.com/2017/09/30/on-leadership-vs-management</id>
  <published>2017-09-30T05:10:59+00:00</published>
  <updated>2017-09-30T05:10:59+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;&lt;em&gt;This post is a flight of ideas. Blame Charity getting my brain going&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This morning, I came to twitter to find &lt;a href=&quot;https://storify.com/dbsmasher/conversation-with-mipsytipsy-and-jaysoifer&quot;&gt;this thread of tweets&lt;/a&gt; from the always awesome Charity majors.&lt;/p&gt;

&lt;p&gt;This got me thinking about how it seems many people in our field, and so many companies by consequence, have adopted the attitude that leadership comes from above. That you must have a certain title to be able to lead in your organization.&lt;/p&gt;

&lt;p&gt;In individual contributors, this can be a sign of lack of maturity. An engineer should not be considered ‘senior’ unless they exhibit a sense of ownership towards the business and a sensibility towards what’s best for the customers. Not only that, but a senior engineer should also be capable of influencing engineers around her to ‘do it right’. That sometimes means the longer way and sometimes even means the less ‘cool’ tech. A senior engineer, even if not a ‘team lead’, should be able to argue, with data, for solutions that provide the most business value with the least amount of risk.&lt;/p&gt;

&lt;p&gt;Similarly, if Junior/associate engineers are to grow in their careers, they should be encouraged to find their passion and translate that into business value. As engineers, we may be insulated from the customer facing work by having support teams and customer success reps but that should not mean we cannot be aware of how the quality of our work impacts the customers and our co workers who have to talk to the customers.&lt;/p&gt;

&lt;p&gt;Sadly we do not learn enough of that in our fancy CS degrees. Too much emphasis is put on algorithms and software and not enough in understanding how to talk to the business side of the company, how to have empathy for the customer facing team members and how to behave and think as one team that is ultimately providing a service to paying customers. I was guilty of this as a fresh grad. I was looking to write code, to play with software that was new to me and handling customers’ issues was a ‘nuisance’ that I just had to deal with. It took the fall of my first company to learn that code and elegant design are nothing if they are not providing a business value and solving a problem for a paying customer.&lt;/p&gt;

&lt;p&gt;This is also a problem of managers. Micro managers, among many other harms, reinforce the idea that only managers can decide how things are done and make decisions. Select all managers who do not advocate for their team and know how to say no during critical organizational planning imply to the team that they cannot drive excellence themselves but have to be told, by some laid out plan devised by executives, what to do.&lt;/p&gt;

&lt;p&gt;Now, that is not to say that the executive perspective is not of value. Far from it, an executive has the position of knowing the market landscape a company is competing in, and owns the strategy of the organization to develop an edge in that market. But the executive is not to be expected to know the details and severity of tech debt the company has accumulated and what parts of said tech debt can closely endanger meeting this needed business edge. Without individual contributors understanding their company strategy and making these connections between “this service is old and needs a refactor” and “we need this to make that new product scalable and an easier sell”, the executive team may never see their lofty plans to fruition and ultimately the business will lose customers.&lt;/p&gt;

&lt;p&gt;So what can be done to make this better?&lt;/p&gt;

&lt;p&gt;Yes it involves everyone…&lt;/p&gt;

&lt;p&gt;It is important to make ‘sense of ownership’ a part of performance review for everyone, not just for team leads and line managers. It should be something junior individual contributors strive to internalize in order to grow in their careers and it needs to be something senior team members have to exhibit and are scored on. Remember, it is what goes in the performance review that shows what the company &lt;em&gt;really&lt;/em&gt; values. Since it is literally ‘putting money where your mouth is’. One thing that is &lt;em&gt;very&lt;/em&gt; dangerous, is mistaking when an individual contributor is sounding the alarm about an architectural problem as them being “a cynic” or a “pessimist”. That is especially a problem that women in tech face as we are expected to just be merry and happy all the time and when we point out issues during design reviews, it tends to be seen as being ‘brash’ or ‘harsh’. Actively burying concerns from team members and chalking them off to ‘personality’ or the ever non inclusive phrase ‘team fit’ will only spell long term dysfunction for your organization. Ignore at your own risk.&lt;/p&gt;

&lt;p&gt;Line managers, those whose direct reports are all individual contributors, need to constantly let their reports bubble up pain points in the company tech stack. Involve them in the roadmap planning process. Make sure to communicate to them the strategy and not just ‘here is what we will be doing’. For many people, knowing the why goes a very long way in being invested to do the best possible job. Line managers should encourage the quiet ones to still participate in this discussion even if not in front of the whole team. Sometimes the best feedback on ‘what needs to be fixed’ can come out of 1:1 conversations. Not everyone is comfortable sounding these concerns in a group.&lt;/p&gt;

&lt;p&gt;Finally, directors and executives should be open to feedback from all levels of the organization. Do not wait for this to come to you. At my current company we do a biannual survey that is deliberately anonymous but allows everyone to provide the executive team with feedback from all levels of the company. It is &lt;em&gt;super&lt;/em&gt; important to not just solicit this feedback but to transparently also create action items based on that feedback and report back to the company on the progress of these action items.&lt;/p&gt;

&lt;p&gt;Thanks to Charity and Nicole for sparking this post and to &lt;a href=&quot;https://www.amazon.com/Managers-Path-Leaders-Navigating-Growth/dp/1491973897/ref=sr_1_1?ie=UTF8&amp;amp;qid=1506747712&amp;amp;sr=8-1&amp;amp;keywords=the+manager&apos;s+path&quot;&gt;Camille Fournier’s book&lt;/a&gt; that has given me a great perspective into management.&lt;/p&gt;


    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2017/09/30/on-leadership-vs-management.html&quot;&gt;On leadership vs management&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on September 30, 2017.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Chef audit mode]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2017/05/07/untitled-chef-audit-mode.html" />
  <id>https://blog.dbsmasher.com/2017/05/07/untitled-chef-audit-mode</id>
  <published>2017-05-07T20:49:11+00:00</published>
  <updated>2017-05-07T20:49:11+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;I have spoken before about how important it is for me and my team to make as many parts of the database stack match our larger infrastructure. One of the most crucial ways to do this is to make sure that not only are we deploying and managing databases using configuraton management, but that the cookbooks remain in lock step with changes our cookbooks at large move towards.&lt;/p&gt;

&lt;p&gt;When i first learned chef, the way to test your chef cookbooks was &lt;a href=&quot;https://github.com/chef/minitest-chef-handler&quot;&gt;chef minitest&lt;/a&gt;. But that is &lt;a href=&quot;https://github.com/chef/minitest-chef-handler/blob/master/README.md#deprecation-notice&quot;&gt;now deprecated&lt;/a&gt; and is no longer the recommended test method of chef cookbooks. So what is a DBA who is not always in tune with chef land to do? learn from her teammates of course! 😀&lt;/p&gt;

&lt;p&gt;With the help of our ops engineering team members, I was directed at some examples in newer cookbooks we have and to the resources supported by InSpec which is what chef-audit is based on.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;So how does chef audit mode work?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Chef audit works exactly like recipe code. In fact, you can write the audit code right inside the recipe it is auditing. It is very natural in its language which makes it very easy to write &lt;em&gt;before&lt;/em&gt; the actual chef code (hint hint: TDD FTW!) and it has lots of &lt;a href=&quot;http://serverspec.org/resource_types.html&quot;&gt;resource types&lt;/a&gt; which make for simple, easy to read and maintain, tests.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;So if I have a cookbook that was using minitest, how can I make it move to chef audit?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You need to decide where your audit code will live. Your options are:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;One gigantic audit recipe that is included in your runlist however you normally include recipes in the cookbook. This option will put all the code in one file but can get unwieldy and large as a cookbook gets more complex&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;An &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;_audit_foo.rb&lt;/code&gt; recipe for every &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;foo.rb&lt;/code&gt; recipe you already have. The audit ones have to be separately included as well. This is arguably the most organized manner to do this and can work very well for books with lots of internal recipes where one large audit file can get very large. But it can also feel like recipe sprawl.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Add the audit code directly inside each recipe. This is nice because you can then see in the same file both the code that makes the changes and the audits that run to validate the policies around these changes. But again, this can get harder to use if the individual recipe is long or complex&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of these, as you can see, have pros and cons. The good thing about this flexibility is that you can pick what works best for your cookbook or organization :D&lt;/p&gt;

&lt;p&gt;Now onto an example…or two…&lt;/p&gt;

&lt;p&gt;Say you have this code block in a recipe to drop a script file in a specific spot and make it executable.&lt;/p&gt;
&lt;div class=&quot;language-ruby highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;cookbook_file&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;/usr/local/bin/test_backup.sh&apos;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do&lt;/span&gt;
  &lt;span class=&quot;n&quot;&gt;source&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;test_backup.sh&apos;&lt;/span&gt;
  &lt;span class=&quot;n&quot;&gt;owner&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;mysql&apos;&lt;/span&gt;
  &lt;span class=&quot;n&quot;&gt;group&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;mysql&apos;&lt;/span&gt;
  &lt;span class=&quot;n&quot;&gt;mode&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;o755&lt;/span&gt;
  &lt;span class=&quot;n&quot;&gt;action&lt;/span&gt; &lt;span class=&quot;ss&quot;&gt;:create&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The code to audit this block would look like this&lt;/p&gt;
&lt;div class=&quot;language-ruby highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;control_group&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;MySQL BackupTests : Archive access&apos;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do&lt;/span&gt;
  &lt;span class=&quot;n&quot;&gt;control&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;test_backup.sh&apos;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do&lt;/span&gt;
      &lt;span class=&quot;n&quot;&gt;describe&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;file&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;/usr/local/bin/test_backup.sh&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;should&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;exist&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;should&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;be_file&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;should&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;be_executable&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;should&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;be_owned_by&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;mysql&apos;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
      &lt;span class=&quot;k&quot;&gt;end&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;end&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;As you can see, the audit part is very natural in its language, and the resources are quite simple to use. So how do we tell chef to run this audit code in our test environment?&lt;/p&gt;

&lt;p&gt;Presuming you use test kitchen, you need to edit that config to enable audit mode.
Under the provisioner section of your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.kitchen.yml&lt;/code&gt; config file, add these 2 lines:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;client_rb:
  audit_mode: :enabled
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;So what are the advantages of audit_mode vs the old minitests? well besides the fact that minitest is deprecated, I find that having the test code as part of the recipe tree (and better yet, can even be &lt;em&gt;in&lt;/em&gt; the recipe directly) gives a very nice single view of what every recipe should look like and what my view of the state of the host should be based on that. That should come even handier for anyone trying to look at a cookbook I wrote and understand what the cookbook is supposed to do.&lt;/p&gt;

&lt;p&gt;Aa I started a journey with chef &lt;a href=&quot;https://dbsmasher.com/2015/02/12/-learning-configuration-management-as-a-dba/&quot;&gt;a few years back&lt;/a&gt;, converting my knowledge on how to build our databases into repeatable cookbooks, I will be spending the next months with the rest of our data ops team converting our resources into chef audit code.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2017/05/07/untitled-chef-audit-mode.html&quot;&gt;Chef audit mode&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on May 07, 2017.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[On being on call]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2016/12/07/on-being-on-call.html" />
  <id>https://blog.dbsmasher.com/2016/12/07/on-being-on-call</id>
  <published>2016-12-07T18:38:21+00:00</published>
  <updated>2016-12-07T18:38:21+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;In case you don’t follow &lt;a href=&quot;https://sysadvent.blogspot.com/&quot;&gt;SysAdvent&lt;/a&gt; (in which case, why?? you are missing out on great content every year!), the post on December 6th this year was by &lt;a href=&quot;https://sysadvent.blogspot.com/2016/12/day-6-no-more-on-call-martyrs.html&quot;&gt;Alice Goldfuss&lt;/a&gt;. I have tweeted a link to that post that day and since then my tweet has been quoted, responded to and the ensuing ‘conversation’ has been eyeopening.&lt;/p&gt;

&lt;p&gt;A lot of the response has been anger that an operations engineer is advocating for putting developers on call. A lot were quick to suggest that this is a quick way to lose engineers, that it advocates a terrible quality of life, that it stomps on work/life balance and that it ultimately doesn’t serve to improve code quality.&lt;/p&gt;

&lt;p&gt;The most surprising response was suggesting that this call comes from the privilege of not having family/kids responsibilities which I suppose implies only the single people are in ops and carrying pagers.&lt;/p&gt;

&lt;p&gt;I can go on and on about how little empathy I have seen from people who are supposed to be fellow engineers towards their fellow human ops people&lt;/p&gt;

&lt;p&gt;I can go on and on about how many naively suggested that testing and ‘QA’ are sufficient to not need the people who wrote the code to own their work in the real world environment of a system of scale&lt;/p&gt;

&lt;p&gt;I can go on and on about what it says about people who presumes Alice meant “give the developer a pager and send off into the night”…I suppose that’s fine when you do it to the ops gal instead.&lt;/p&gt;

&lt;p&gt;There is a lot to unpack in these responses, many of which disappointing from engineers I respect, but I want to pinpoint 2 things&lt;/p&gt;

&lt;p&gt;What does it mean to put engineers on call&lt;/p&gt;

&lt;p&gt;NO it does not mean spite and no it is not punishment for bugs and no it is not a call for dissolving the lines between life and work. I have written about my experience with burnout before and I wish that on no one.&lt;/p&gt;

&lt;p&gt;What it means to put engineers who write the code on call is that they are the subject matter experts on what is running in production. They know when error foo happens it means the cache layer failed, DNS is taking too long or maybe that the other microservice in that product comprised of about a dozen of them is the one silently failing. You may think an architecture diagram can make stuff like this clear as day but when YOU, the person on the team who has been building this thing for 2 quarters has to look at your own architecture diagram one quarter later to grok which microservice has failed you will realize why you should be paged first.&lt;/p&gt;

&lt;p&gt;This does not mean at all that Ops wants nothing to do with your application paging out at 3 AM. We put our delivery engineering teams (the term ‘delivery’ here is deliberate and very appropriate) on call but they always have an escalation path to ops that is not gated by a timer. Ops still has your back. If you think this is a network problem you are not familiar with, you can immediately page the on call ops engineer and she will help verify where the broken zeros and ones are.&lt;/p&gt;

&lt;p&gt;This is not an effort to spite those who dare push a bug to production. Anyone who has been responsible for production environments knows that will happen. Hell, you don’t even HAVE to introduce bugs through a deploy, really. Unforeseen changes in your environment will cause presumptions to stop being valid and take both those who wrote the code and those with less knowledge of it (Ops) by surprise.&lt;/p&gt;

&lt;p&gt;What this does mean is that we are explicitly saying that the code is useless unless it provides the customers the value they paid for. And when a C level executive has to explain to customers why an outage took X time to resolve, “we couldn’t get the engineering team involved for some time” is not an acceptable answer.&lt;/p&gt;

&lt;p&gt;Your code and my servers and databases are a risk center unless we are both invested in making them run. It is as simple as that.&lt;/p&gt;

&lt;p&gt;What does it mean to be a senior engineer&lt;/p&gt;

&lt;p&gt;It honestly frightened me that people who are senior engineers are balking so hard at giving developers pagers. Maybe they fell for the hyperbole of “we now don’t want developers to sleep either”, maybe they’ve been on call in difficult environments before and that’s the PTSD talking. But I do not believe that any engineer should be called ‘senior’, by title or by implication if they refuse to be reachable in a team rotation in case their own work caused a customer facing issue.&lt;/p&gt;

&lt;p&gt;Note my words because I am trying my best to choose them carefully. ‘customer facing issue’. You can’t give people pagers without making sure you are not paging on bullshit signals that aren’t actually affecting the customers and the business. And surprise, ops engineers are people too.&lt;/p&gt;

&lt;p&gt;So what does this mean? If you fail to see how your code is more than the sum of its functions and test. That it needs to provide real value. If you insist that someone else be the first line of defense when it breaks, then you are failing to acknowledge that as a senior engineer, your job is more than producing code. You are setting a terrible example for junior engineers on your team. They will now learn the lesson of “I don’t have to own that my code provides value once it’s in production”. The damage that mentality will do to an engineering organization is lasting and will take a long time before anyone realizes it is the root for a LOT of tech debt and it takes years to also reverse.&lt;/p&gt;

&lt;p&gt;If you are a manager promoting engineers to senior titles and an emphasis on ownership, including stability, is not a non-negotiable criterion of that promotion then you are damaging your organization. And if you think that a sense of ownership and truly understanding what it takes to create stable systems of scale can happen without ever being paged by a service in production then I am not sure why you are in charge of engineers supposedly building such systems of scale.&lt;/p&gt;

&lt;p&gt;Finally, no one is saying this all ignores our lives outside work. In fact, the opposite. Ops engineers have lives too. We are men and women with spouses and babies who get sick and sometimes are solo parenting. Mature teams have each others’ backs. Mature teams will override set on call shifts when life strikes and that is applauded.&lt;/p&gt;

&lt;p&gt;What I (and I think Alice) are saying is “do not systematically make it acceptable to throw code at the production wall”….not let’s page developers at 3 AM out of spite.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2016/12/07/on-being-on-call.html&quot;&gt;On being on call&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on December 07, 2016.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Capacity planning for databases]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2016/03/25/capacity-planning-databases.html" />
  <id>https://blog.dbsmasher.com/2016/03/25/capacity-planning-databases</id>
  <published>2016-03-25T03:17:11+00:00</published>
  <updated>2016-03-25T03:17:11+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;&lt;em&gt;Note:&lt;/em&gt; This is inspired by Julia Evans’ recent post about ….&lt;a href=&quot;http://jvns.ca/blog/2016/03/20/how-do-you-do-capacity-planning/&quot;&gt;capacity planning&lt;/a&gt; 😌&lt;/p&gt;

&lt;h3 id=&quot;ground-rules&quot;&gt;Ground rules&lt;/h3&gt;

&lt;h4 id=&quot;rdbms&quot;&gt;RDBMS&lt;/h4&gt;
&lt;p&gt;Yes..this post is geared for those of us who use MySQL with a single writer at a time and 2 or more read replicas. A lot of what I will talk about here applies differently, or not at all, to multi writer clustered datastores, although those also come with their own set of compromises and caveats. So…your milage will definitely vary.&lt;/p&gt;

&lt;h4 id=&quot;sharding&quot;&gt;Sharding&lt;/h4&gt;

&lt;p&gt;I have already covered large strokes of this in one of my &lt;a href=&quot;https://blog.dbsmasher.com/2015/02/08/scaling-mysql-at-sendgrid/&quot;&gt;earlier posts&lt;/a&gt;, I mostly focused there on the benefits of functional or horizontal sharding. Yes that is a prerequisite, since what you use to access the database layer WILL decide how much flexibility you have to scale.&lt;/p&gt;

&lt;p&gt;If you are a company that experiences large differences between peak and average traffic, you should be prepared to leave the paradigm of ‘the database’ as a single physical entity behind.&lt;/p&gt;

&lt;h4 id=&quot;ability-to-split-reads-and-writes&quot;&gt;Ability to split reads and writes&lt;/h4&gt;

&lt;p&gt;This is something you will need to be able to do, but not necessarily enforce as a set in stone rule. There will be use cases where a write needs to be read very soon after and where tolerance for things like lag/eventual consistency is low. Those are ok to have, but in the same applications, you will also have scenarios for reads that &lt;em&gt;can&lt;/em&gt; tolerate some longer time span of eventual consistency. When such reads are in high volume, do you really want that volume going to your single writer if it doesn’t really have to? Do yourself a favor, and make sure soon in your growth days that you can control the use of a read or write IP in your code.&lt;/p&gt;

&lt;p&gt;Now onto the thought process of actual capacity planning…&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A database cluster is not keeping up. what do I do?&lt;/strong&gt;&lt;/p&gt;

&lt;h4 id=&quot;determine-the-system-bottleneck&quot;&gt;Determine the system bottleneck&lt;/h4&gt;

&lt;ul&gt;
  &lt;li&gt;Is the issue high CPU?&lt;/li&gt;
  &lt;li&gt;Is it IO capacity?&lt;/li&gt;
  &lt;li&gt;Is it growing lag without a clear query culprit?&lt;/li&gt;
  &lt;li&gt;Is is locks?&lt;/li&gt;
  &lt;li&gt;How do I know which it is?&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id=&quot;you-need-a-baseline&quot;&gt;You need a baseline&lt;/h4&gt;
&lt;p&gt;Once you know what system metric you are mostly bound to, you need to establish baseline and peak values. Otherwise, determining whether your current issue is a bug vs real growth is going to be a lot more error prone than you’d like.&lt;/p&gt;

&lt;p&gt;Basic server metrics can only go so far but at some point you will find you also need context based metrics. Query performance and app side perceived performance will tell you what the application sees as a response time to queries.&lt;/p&gt;

&lt;h4 id=&quot;learn-your-business-traffic-patterns&quot;&gt;Learn your business traffic patterns&lt;/h4&gt;
&lt;p&gt;Are you a business that is susceptible to peaks in specific weekdays (marketing)? do you have regular launches that triple or quadruple your traffic like gaming? These sorts of questions will drive how much of reserved headroom you should keep or whether you need to invest in elastic growth.&lt;/p&gt;

&lt;h4 id=&quot;determine-the-ratio-of-raw-traffic-numbers-in-relation-to-capacity-in-use&quot;&gt;Determine the ratio of raw traffic numbers in relation to capacity in use&lt;/h4&gt;
&lt;p&gt;This is simply the answer to “If we made no code optimizations, how many emails/sales/whatever” can we serve with the database instance we have right now?&lt;/p&gt;

&lt;p&gt;Ideally, this a specific value that makes the math towards planning a year’s growth a simple math equation. But life is never ideal and this value will vary depending on season or completely external happy factors like signing up a new major customer. In early startups this number is a faster moving target but it should stabilize as the company transitions from early days to more established business with more predictable business growth patterns.&lt;/p&gt;

&lt;h4 id=&quot;do-i-really-need-to-buy-more-machines&quot;&gt;Do I really need to buy more machines?&lt;/h4&gt;
&lt;p&gt;You need to find a way to determine if this is truly capacity (I need to split the writes to support more concurrent write load or add more read replicas) vs code based performance bottleneck (this new query can really have its results cached in something cheaper and not beat the database as much).&lt;/p&gt;

&lt;p&gt;How do you do that? You need to get familiar with your queries. The baby step for that is a combination of &lt;a href=&quot;http://innotop.googlecode.com/svn/html/manual.html&quot;&gt;innotop&lt;/a&gt;, slow log and the &lt;a href=&quot;https://www.percona.com/software/mysql-tools/percona-toolkit&quot;&gt;percona toolkit&lt;/a&gt;’s pt-query-digest. You can automate this by shipping the DB logs to a central location and automating the digest portion.&lt;/p&gt;

&lt;p&gt;But that is also not the entire picture, slow logs are performance intensive if you lower their threshold too much. If you need less selective sampling you will need to detect the entire conversations between the application and the datastore. In open source land you can go as basic as tcpdump or you can use hosted products like datadog, newrelic or vivid cortex.&lt;/p&gt;

&lt;h4 id=&quot;make-a-call&quot;&gt;Make a call&lt;/h4&gt;
&lt;p&gt;Capacity planning can be 90% science and 10% art but that 10% shouldn’t mean that we shouldn’t strive for as much of the picture as we can. As engineers we can sometimes fixate on the missing 10% and not realize that if we did the work, that 90% can get us v far into a better idea of our stack’s health, a more efficient use of our time optimizing performance and planning capacity increases carefully which eventually results in much better return on investment for our products.&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2016/03/25/capacity-planning-databases.html&quot;&gt;Capacity planning for databases&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on March 25, 2016.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[2015 in review...2016 here I come]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2016/01/04/2015-review.html" />
  <id>https://blog.dbsmasher.com/2016/01/04/2015-review</id>
  <published>2016-01-04T18:22:21+00:00</published>
  <updated>2016-01-04T18:22:21+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;p&gt;&lt;em&gt;This post is inspired by quite a few great posts like this one by &lt;a href=&quot;http://beero.ps/2016/01/01/on-to-2016/&quot;&gt;Ryn Daniels&lt;/a&gt; and &lt;a href=&quot;http://larahogan.me/blog/2014-2015/&quot;&gt;Lara Hogan&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;2015 was a year of lots of changes and lots of growing for me, both professionally and personally. It also had a lot of firsts for me.&lt;/p&gt;

&lt;h3 id=&quot;going-remote&quot;&gt;Going remote&lt;/h3&gt;

&lt;p&gt;Late in 2014, I announced to Sendgrid that I am about to move to South Florida where my husband has been working for months. Sendgrid, while distributed in multiple offices, always had all its employees in one of these offices, no truly remote employees. So this was new territory and it was a new challenge given how I am usually collaborating with more than just the Operations team.&lt;/p&gt;

&lt;p&gt;After a year, while I do not declare it an absolute success, i feel that I would not still be working at Sendgrid if it weren’t for my team’s and managers’ support. It is easy to declare someone a ‘good employee’, with all the different meanings different companies embue that term with. But the support I got from the team was absolutely crucial to my ability to continue providing value to the company.&lt;/p&gt;

&lt;h3 id=&quot;first-tech-conference&quot;&gt;First Tech conference&lt;/h3&gt;

&lt;p&gt;People still get surprised when they hear this but until June of 2015, I had not been to a single tech conference. A combination of being permanently busy, being a mom, moving across country more than a couple of times
have led to having no time to properly plan any tech conference attendance. Finally I took the plunge and I picked &lt;a href=&quot;http://monitorama.com/&quot;&gt;Monitorama&lt;/a&gt; as my first one.&lt;/p&gt;

&lt;p&gt;The talks were great. One talk I couldn’t stop referring to since returning was &lt;a href=&quot;https://vimeo.com/131484322&quot;&gt;“Engineering happiness”&lt;/a&gt; By Laura Thompson of Mozilla. I could see myself, post and present co workers in the examples she provided on how a happy engineer slowly becomes less happy with the status quo. One of the first things I did once I returned to work, was share that video with my team lead and engineering leader at the time. I felt it was important to share her advice and to internalize it.&lt;/p&gt;

&lt;p&gt;However, conferences are not just about the talks. While I learned a lot in those, I found the hallway track the real gem of that trip and the reason that conference has left lasting memories with me.&lt;/p&gt;

&lt;p&gt;I got to meet lots of people I had been following on twitter for a while that I already learned a lot from, and looked up to. I had a LOT of fun talking about being a woman in the ops community with &lt;a href=&quot;https://twitter.com/beerops&quot;&gt;Katherine Daniels&lt;/a&gt; and &lt;a href=&quot;https://twitter.com/sigje&quot;&gt;Jennifer Davis&lt;/a&gt;. Who are about to publish a &lt;a href=&quot;http://shop.oreilly.com/product/0636920039846.do&quot;&gt;book&lt;/a&gt; I am looking forward to. Another set of conversations that were very educational and have pushed some of my 2016 plans was talking about team leadership and management challenges with Roy Rapaport of Netflix.&lt;/p&gt;

&lt;p&gt;There are too many people to list here that I enjoyed meeting and talking to at Monitorama. And a lot of those conversations shaped and guided things I have done the rest of the year and things I have planned for 2016.
I cannot recommend enough how great this conference was. Hat tip to &lt;a href=&quot;https://twitter.com/obfuscurity&quot;&gt;Jason Dixon&lt;/a&gt; for creating that great environment and all the #hugops&lt;/p&gt;

&lt;h3 id=&quot;blog-and-blog-posts&quot;&gt;Blog and blog posts&lt;/h3&gt;

&lt;p&gt;I started this blog early in 2015. It was not the first time I start a blog and I was not sure if I was gonna keep it up. But then I wrote a blog post for my company and I decided that maybe I will want to write more.&lt;/p&gt;

&lt;p&gt;Yes part of that is building personal brand but I was also sometimes frustrated by talking about important things in Devops or in database architecture in 140 character pieces. Yes I did start the blog early in 2015 but it was a blog post by Lara Hogan about &lt;a href=&quot;http://larahogan.me/blog/celebrate-achievements/&quot;&gt;celebrating our achievements&lt;/a&gt;, conversations with my boss in late 2014 and a personal conversation with Jennifer Davis that convinced me that I do have things to say and that my experience so far at Sendgrid could be beneficial to others who are just starting as DBAs.&lt;/p&gt;

&lt;p&gt;Most of what I wrote this year was technical or ‘lessons learned’ in the roller coaster of managing databases at a fast growing company. But I also wrote one post that is near and dear to my heard about &lt;a href=&quot;https://blog.dbsmasher.com/2015/07/29/on-burnout/&quot;&gt;burnout&lt;/a&gt; that got me to face how much I needed to take care of me at the same level of commitment I was taking care of servers and databases.&lt;/p&gt;

&lt;h3 id=&quot;2016-plans&quot;&gt;2016 Plans&lt;/h3&gt;
&lt;ul&gt;
  &lt;li&gt;Give a tech conference talk. I submitted my &lt;a href=&quot;https://www.percona.com/live/data-performance-conference-2016/sessions/bringing-devops-dbas-chef&quot;&gt;first ever abstract &lt;/a&gt;. Don’t know yet if it will be accepted but I am excited to go to PLCME for the first time nonetheless&lt;/li&gt;
  &lt;li&gt;Take better care of me. We hired our second DBA at Sendgrid late 2015. The signs of burnout on me were clear. Nothing is clearer when one performance review says “I fear that Silvia works too hard. She needs to take care of herself” 😊. With onboarding the new member, I hope to be able to split the never ending list of things we need to get done in 2016&lt;/li&gt;
  &lt;li&gt;A deeper focus on architecture and automation. I have spent a huge portion of the last few years working with engineers on schema design and bringing the first step of database configuration management to fruition. But a mature infrastructure is more than just configuration management and I hope to be able to grow more skills in larger system design, making the database layer more robust and a true PaaS layer&lt;/li&gt;
  &lt;li&gt;Write more technical blog posts. We do so much at Sendgrid. I think we can share lots of lessons learned and I hope I can help with that.&lt;/li&gt;
  &lt;li&gt;Get closer to becoming Staff Engineer.&lt;/li&gt;
  &lt;li&gt;Much more…&lt;/li&gt;
&lt;/ul&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2016/01/04/2015-review.html&quot;&gt;2015 in review...2016 here I come&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on January 04, 2016.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Using Sensu for DBA tasks]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2015/11/03/using-sensu-for-dba-tasks.html" />
  <id>https://blog.dbsmasher.com/2015/11/03/using-sensu-for-dba-tasks</id>
  <published>2015-11-03T21:31:47+00:00</published>
  <updated>2015-11-03T21:31:47+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;h3 id=&quot;sensu-for-monitoring&quot;&gt;Sensu for monitoring&lt;/h3&gt;
&lt;p&gt;Here at Sendgrid we spent the last couple of years porting a lot of our service and host monitoring to &lt;a href=&quot;https://sensuapp.org/&quot;&gt;Sensu&lt;/a&gt;. Its solid API support meant we could write all sorts of tooling around it. We also liked the idea of standalone, client side checks that push their status to the Sensu alerting queue asynchronously. If you are new to Sensu or haven’t ever read on it, &lt;a href=&quot;https://sensuapp.org/docs/latest/overview&quot;&gt;this is a good place to start&lt;/a&gt;.&lt;/p&gt;

&lt;h3 id=&quot;typical-usage-example&quot;&gt;Typical usage example&lt;/h3&gt;

&lt;p&gt;The typical use I have for such standalone checks is health checks, a simple example looks like this in Sensu’s client config&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;{
  &quot;checks&quot;: {
    &quot;mysql_alive&quot;: {
      &quot;command&quot;: &quot;mysql-alive.rb -h &amp;lt;IP&amp;gt; -d mysql -u :::mysql.user|sensu::: -p :::mysql.password:::&quot;,
      &quot;handlers&quot;: [
        &quot;default&quot;
      ],
      &quot;standalone&quot;: true,
      &quot;interval&quot;: 10,
      &quot;notification&quot;: &quot;OMG MySQL is dead!&quot;
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;But if you look closer, all you really do is tell Sensu to run a command. So this can be…any command. This will be useful in a just a moment.&lt;/p&gt;

&lt;h3 id=&quot;what-i-am-solving&quot;&gt;What I am solving&lt;/h3&gt;

&lt;p&gt;I have been traditionally using the crond service for running local management tasks on databases like rotating partitions and triggering backups. Along with that, we have a report that would check the logs of these jobs, on each backup replica, and make sure they ended in success messages. This is fragile due to a number of reasons:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;CronD doesn’t have any built in monitoring. It does not have a concept of ‘stale’&lt;/li&gt;
  &lt;li&gt;Those reports - I will call them ‘watchers’ - are one more moving part adding complexity to the question ‘when was the last successful backup of my DB?’&lt;/li&gt;
  &lt;li&gt;This setup is prone to race conditions. You must time the task and its watcher in cron exactly or else the watcher can preemptively trigger an alert or signal failure when the backup is not done yet. Any drift in duration of the task will eventually make this happen (like a backup taking longer as a database grows).&lt;/li&gt;
  &lt;li&gt;What if the watcher script didn’t run? It is also in CronD - either on the same host or on another host, right? Either you find yourself in a rabbit hole of who watches the watchers, or a human has to notice that a report didn’t come out…humans aren’t good at remembering things.&lt;/li&gt;
  &lt;li&gt;Changing the designation of a server means you must change it in a number of places or the watcher will watch the wrong host.&lt;/li&gt;
  &lt;li&gt;We are striving to &lt;a href=&quot;http://mcfunley.com/choose-boring-technology&quot;&gt;keep our stack boring&lt;/a&gt;. While newer technologies like chronos and rundeck provide more enhanced scheduling, they also need a service discovery layer to do this right. That was a bigger undertaking and too much scope creep for what I was solving.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So I decided to make Sensu work to my advantage, with the help of chef roles.&lt;/p&gt;

&lt;h3 id=&quot;partition-rotation-in-sensu&quot;&gt;Partition rotation in Sensu&lt;/h3&gt;

&lt;p&gt;I started off by moving any credentials I need for my partition rotation script into Sensu’s redacted configuration. This is good practice for &lt;em&gt;anything&lt;/em&gt; you put into sensu that uses secrets. The credentials are added in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/etc/sensu/client.json&lt;/code&gt; then are just referenced in check configurations using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;:::secret_thingie:::&lt;/code&gt; notation.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;_Protip_&lt;/em&gt;: Sensu doesn’t know what to redact in client.json. You must also define the name of the keys you want redacted. like this..&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&quot;redact&quot;: [
  &quot;other_password&quot;,
  &quot;password&quot;
]
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;This is also not a deep merged list. As you can see I had to explicitly include ‘password’ once I needed to add another key in that list.&lt;/p&gt;

&lt;p&gt;Then I needed to define the new Sensu check that rotates the partitions. I am using a chef resource as the code example.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;sensu_check &apos;add_table_partitions&apos; do
  command &quot;/usr/local/pdb/bin/pdb-parted --add --interval d +7d.startof h=localhost,u=specialdbuser,p=:::redacted_partition_password:::,D=mydb,t=special_table &amp;gt;&amp;gt; /var/log/partition_rotation.log 2&amp;gt;&amp;amp;1&quot;
  handlers %w(default_handler special_dba_handler)
  interval 86400
  standalone true
  additional(:occurrences =&amp;gt; 3, :notification =&amp;gt; &quot;#{node[&apos;hostname&apos;]} failed adding partitions to important table. See log file in /var/log/mail_send_cancel_pause.log&quot;)
end
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Let’s inspect what just happened here…&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command&lt;/code&gt;:  &lt;a href=&quot;https://github.com/palominodb/PalominoDB-Public-Code-Repository/blob/master/tools/data_mgmt/t/pdb-parted/pdb-parted.t&quot;&gt;pdb-parted&lt;/a&gt; is a very useful perl script by the folks from Palominodb (now at &lt;a href=&quot;http://www.pythian.com/&quot;&gt;Pythian&lt;/a&gt;) for rotating partitions in a MySQL DB. This is the same line I used to maintain in a crontab file configuration.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;interval&lt;/code&gt;: how often this Sensu check runs. In this example these are daily partitions so running the script daily was sufficient. &lt;em&gt;The important part here is that your script is idempotent&lt;/em&gt;. pdb-parted is. If it finds that the needed partitions already exists, it just outputs a message to that effect and exits nicely.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;occurrences&lt;/code&gt;: This is the number of allowed failures before alerting. It is nice to have a buffer especially when I already configure the command to make partitions days/weeks in advance.&lt;/p&gt;

&lt;p&gt;For those who use tools other than chef and want to see what the final check configuration looks like, here it is, important bits redacted:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;{
  &quot;checks&quot;: {
    &quot;add_table_partitions&quot;: {
      &quot;command&quot;: &quot;/usr/local/pdb/bin/pdb-parted --add --interval m +7m.startof h=localhost,u=specialdbuser,p=:::redacted_partition_password:::,D=mydb,t=special_table &amp;gt;&amp;gt; /var/log/partition_rotation.log 2&amp;gt;&amp;amp;1&quot;,
      &quot;handlers&quot;: [
        &quot;default_handler&quot;,
        &quot;special_dba_handler&quot;
      ],
      &quot;standalone&quot;: true,
      &quot;interval&quot;: 86400,
      &quot;occurrences&quot;: 3,
      &quot;notification&quot;: &quot;hostname failed adding partitions to special table. See log file in /var/log/partition_rotation.log&quot;
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Partition rotation now has built in monitoring. If this script exits with a non zero code, the Sensu handlers take care of getting this knowledge to the predefined group of people using the method we want. We already have multiple on-call team rotations set up and escalation paths and all the other stuff that comes along with monitoring at scale. So we got to just leverage all of that rather than reinventing the wheel.&lt;/p&gt;

&lt;h3 id=&quot;planned-future-improvements&quot;&gt;Planned future improvements&lt;/h3&gt;

&lt;p&gt;No solution is perfect. Neither is this one.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;We still run the risk of more than one machine in a cluster holding the primary role. Until service discovery is in place, this is something we can build a check for leveraging chef search.&lt;/li&gt;
  &lt;li&gt;Now that tasks like backups are running in a first class citizen tool in our stack, I can more easily add stats and get graphs for how long backups take, how big backups are to make capacity planning and MTTR tracking easier.&lt;/li&gt;
&lt;/ul&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2015/11/03/using-sensu-for-dba-tasks.html&quot;&gt;Using Sensu for DBA tasks&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on November 03, 2015.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[On Burnout]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2015/07/29/on-burnout.html" />
  <id>https://blog.dbsmasher.com/2015/07/29/on-burnout</id>
  <published>2015-07-29T20:23:02+00:00</published>
  <updated>2015-07-29T20:23:02+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;h3 id=&quot;the-not-so-merry-go-round&quot;&gt;The not-so-merry go round&lt;/h3&gt;

&lt;p&gt;There is a lot of talk these days about burnout in our field. A lot of &lt;a href=&quot;http://burnout.io&quot;&gt;great&lt;/a&gt; &lt;a href=&quot;http://mhprompt.org&quot;&gt;initiatives&lt;/a&gt; to get us tech people to not hide it, not sweep it under the rug.&lt;/p&gt;

&lt;p&gt;This is a great start for a more honest conversation about the stresses we all deal with in this industry and yet, a lot of the examples I see of people trying to deal with it is through quitting their current gig, taking a long vacation, then working on finding the next gig.&lt;/p&gt;

&lt;p&gt;In the long run, this is not great for companies’ overall health and more importantly for a team’s health and morale. There are already statistics about how &lt;a href=&quot;http://www.techrepublic.com/blog/career-management/tech-companies-have-highest-turnover-rate/&quot;&gt;churn in IT is the highest compared to other fields&lt;/a&gt; and if you are at all familiar with all the effort that goes into recruiting, hiring and on boarding in our field, you probably already know that it is an expensive process with a lead time in the months. A lot of what i ruminated over the last few weeks wasn’t just my words but also the words of a &lt;a href=&quot;http://logikal.is/blog/2013/04/30/a-month-to-myself/&quot;&gt;dear colleague&lt;/a&gt; who had to do this exact thing a few years back.&lt;/p&gt;

&lt;h3 id=&quot;seeing-it&quot;&gt;Seeing it&lt;/h3&gt;

&lt;p&gt;It may be irrelevant whether our industry makes us wanna appear invincible or if this industry just attracts those of us who want to always appear strong and have it all figured out. But I know &lt;em&gt;I&lt;/em&gt; always try to do that. Family, kid, plus managing a database layer that has expanded twenty-fold in my three and a half years with my current job.&lt;/p&gt;

&lt;p&gt;Let me start with saying that I do not do it alone. Early on we hired consultants to help off load a lot of the DBA tasks. But in the end of the day, I’ve always felt that I am the in house DBA, I own the performance, management, and health of these systems.&lt;/p&gt;

&lt;p&gt;This sense of ownership, while lauded, led me to unrealistic expectations of myself. I was checking work chat all the time, checking email all the time. I had slipped into a pattern of coupling being online/available all the time with doing a good job.&lt;/p&gt;

&lt;p&gt;As the months and quarters rolled by, cracks started to appear.&lt;/p&gt;

&lt;p&gt;I was having less fun with spending time with developers, talking distributed system architecture. Suddenly, I was having the ‘Sunday evening dread’…something I didn’t really think would happen working on so many exciting things and interesting problems for so long.  I could see the snark levels increase in my conversations..&lt;/p&gt;

&lt;p&gt;“yeah I will test this new shiny thing in maybe a few years”&lt;/p&gt;

&lt;p&gt;“Sure…we will someday deprecate this old thing!”&lt;/p&gt;

&lt;h3 id=&quot;corroborating-evidence&quot;&gt;Corroborating evidence&lt;/h3&gt;

&lt;p&gt;Sometimes even though we know something, we need an external source with more experience to confirm to us that it is true, that we aren’t just ‘not good enough’. For me, a lot of that was talks by &lt;a href=&quot;https://twitter.com/lxt&quot;&gt;Laura Thompson&lt;/a&gt; of Mozilla. The &lt;a href=&quot;http://original.livestream.com/etsycodeascraft/video?clipId=pla_b7da40fe-ac51-4a87-ac0b-6002899457eb&amp;amp;utm_source=lslibrary&amp;amp;utm_medium=ui-thumb&quot;&gt;first&lt;/a&gt; is more directed at managers but it absolutely helped me. The &lt;a href=&quot;https://vimeo.com/131484322&quot;&gt;second&lt;/a&gt; was at this year’s Monitorama. All the signs were there. The decrease in my github activity, the constant feeling that I was fighting emergencies all the time, physical and emotional exhaustion, sense of ineffectiveness and lack of accomplishment (Yes, I am now directly quoting the slides)&lt;/p&gt;

&lt;h3 id=&quot;what-to-do&quot;&gt;What to do&lt;/h3&gt;

&lt;h4 id=&quot;own-your-boundaries&quot;&gt;Own your boundaries&lt;/h4&gt;

&lt;p&gt;Being able to separate the time you are working from the time you are not is paramount. In the end, no one owns my well being more than me. I work remote in a different timezone than my team so boundaries were extra important to establish.&lt;/p&gt;

&lt;p&gt;I started off by disallowing any push notifications for work hipchat on my phone. Not accepting meetings past 5 PM. Removing work email from the phone and most importantly, the laptop doesn’t leave the office room during the week.&lt;/p&gt;

&lt;p&gt;What helped me the most was turning off work email on my phone. I needed to accept that email is an asynchronous method of communication and that I shouldn’t feel guilty about not checking it every hour including right before going to sleep. I know this may sound incredibly obvious to some but for me the checking work email and work chat from the phone constantly was like an itch and it took me a while to accept that it was an expectation I was setting for myself and that in the long run it was not making me a more productive employee or a better engineer in any way.&lt;/p&gt;

&lt;h4 id=&quot;talk-to-your-team&quot;&gt;Talk to your team&lt;/h4&gt;

&lt;p&gt;The roles here differ depending on whether you are a manager, or individual contributor. None of this could work without support from my team, including management.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Managers, this is how you avoid churn&lt;/em&gt;.  You need to make sure your team feels safe saying they need a break and to feel safe taking it. None of this ‘unlimited vacation time’ nonsense if no one is actually going on vacation. 1:1s are supremely important here. I am not ignoring that a busy team also means an overwhelmed manager sometimes but this is the time where you as a leader must prioritize keeping in touch with the team than anything else.&lt;/p&gt;

&lt;p&gt;/soapbox&lt;/p&gt;

&lt;p&gt;I am an Individual contributor but I am also one of the more senior team members (in tenure) so accepting that I carry some of the responsibility of setting the tone was necessary. Besides being honest with myself, I needed to be honest with my team. I started letting my project manager and my lead know that I would be staying off chat in the evening. Making sure they have a way to get a hold of me and trusting &lt;em&gt;them&lt;/em&gt; that they will truly only use it sparingly and in emergencies. Without that framework and their support, I may know what I need but I would not feel empowered to act on it and I am very grateful that they let me do that and do the same for themselves so we can all continue working together.&lt;/p&gt;

&lt;h4 id=&quot;talk-to-someone&quot;&gt;Talk to someone&lt;/h4&gt;

&lt;p&gt;This doesn’t have to be a mental health professional although that is also a good thing to do. But in the simpler sense it helps a lot to talk to people who have been or are still going through the same thing, even with a few minor details altered. There is a lot of us in &lt;a href=&quot;http://signup.hangops.com&quot;&gt;hangops.slack.com&lt;/a&gt; who have stories and scars from this. So much so that we have a dedicated mental health channel.&lt;/p&gt;

&lt;h4 id=&quot;work-in-progress&quot;&gt;Work in progress&lt;/h4&gt;

&lt;p&gt;I mentioned in the beginning how I see most people deal with this situation. And there many of us. Does this mean I am looking for a new gig? No. This is not a quitting post :) I like my team. A lot. And I don’t just want to continue working with them but to continue to enjoy it and I want to see them also have a good time working with me. This work in progress has to always start with me recognising what I need and communicating it but I also have a team that has supported the steps I took to ease the stress.&lt;/p&gt;

&lt;p&gt;I can’t stress this enough. I am still learning how to deal with this. I know now I always will be. It is a fine line between being someone who is honest about their work, always caring to go the extra mile and do right by their employer without sacrificing their sanity and inner peace.&lt;/p&gt;

&lt;p&gt;I don’t want to eventually feel contempt and resentment towards that work. I am also not arguing for slacking off in the name of life work balance. Too many companies don’t put a lot of value in keeping their employees healthy in mind as well as body and that is a very good reason to look for another gig, and there are companies that are aware of these pressures and challenges but are only beginning to truly acknowledge them and begin a conversation about them.&lt;/p&gt;

&lt;p&gt;It is on all of us to not try to be individual heroes/ninjas/rockstars and instead promote teams of healthy, rested, smart engineers with well balanced lives.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Many thanks to Sean Kilgore, Jennifer Davis and Charity Majors for helping put these thoughts to words&lt;/em&gt;&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2015/07/29/on-burnout.html&quot;&gt;On Burnout&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on July 29, 2015.&lt;/p&gt;
  </content>
</entry>


<entry>
  <title type="html"><![CDATA[Alfred, csshx and terminalception]]></title>
 <link rel="alternate" type="text/html" href="https://blog.dbsmasher.com/2015/03/06/alfred-csshx-and-terminalception.html" />
  <id>https://blog.dbsmasher.com/2015/03/06/alfred-csshx-and-terminalception</id>
  <published>2015-03-06T14:47:41+00:00</published>
  <updated>2015-03-06T14:47:41+00:00</updated>
  <author>
    <name></name>
    <uri>https://blog.dbsmasher.com</uri>
    <email></email>
  </author>
  <content type="html">
    &lt;h3 id=&quot;alfred-csshx-and-terminalception&quot;&gt;Alfred, csshx and terminalception&lt;/h3&gt;

&lt;p&gt;I use Tmux usually but Tmux on the mac has not been playing nice with csshx for me. Something in the dark magic of perl broke with an error that looks like this&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Mar  5 08:58:42 silvias-MacBookPro.local perl[10828] &lt;Error&gt;: ImageIO: CGImageDestinationFinalize image destination must have at least one image
2015-03-05 08:58:42.473 perl5.18[10828:2436454] CGImageDestinationFinalize failed for output type &apos;public.tiff&apos;
**** ERROR **** PerlObjCBridge:: convertPerlToObjC(): Referenced thingy not blessed
**** ERROR **** PerlObjCBridge:: convertArg() for index 2: convertPerlToObjC() failed
**** ERROR **** PerlObjCBridge:: sendObjcMessage: Error converting argument 1 for message &quot;setObject:forKey:&quot;
**** ERROR **** PerlObjCBridge: error [1] sending message [__NSDictionaryM setObject:forKey:] at /System/Library/Perl/Extras/5.18/darwin-thread-multi-2level/PerlObjCBridge.pm line 248.&lt;/Error&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But I wasn’t about to let that take away my terminalception magic :)&lt;/p&gt;

&lt;p&gt;Note: This is for a mac environment but you may emulate it in your favourite distro by replacing Alfred with whatever launcher you may use.&lt;/p&gt;

&lt;h4 id=&quot;what-you-need&quot;&gt;What you need&lt;/h4&gt;
&lt;p&gt;Alfred app is my favorite launcher in mac but I suspect Quicksilver (if you are the quaint type) can also run commands directly to the terminal.&lt;/p&gt;

&lt;p&gt;Next install Homebrew. You need this to install csshx easily. Also because GNU tools :)
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ruby -e &quot;$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/master/install)&quot;&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Install csshx&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;brew install csshx&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;If you are chef shop like mine, a lot of times what you are looking for is to ssh to all the machines of a certain type. In chef, we call those roles. This step will change depending on the configuration management/service discovery framework of choice in your infrastructure.&lt;/p&gt;

&lt;p&gt;First, set the terminal app your alfred will use. I set it to Terminal because I wanted this to be separate from my usual workspace in iTerm 2&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://farm1.staticflickr.com/735/22396914786_e46c0ed4cd_b.jpg&quot; alt=&quot;Alfred settings&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Then, when I want to fire up csshX to a bunch of our our servers all at once, it usually looks like this.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Protip&lt;/em&gt;: knife ssh can take cssh as an argument, so no awk and bash pipes required.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://farm6.staticflickr.com/5793/22433815321_7b1f1fced2_b.jpg&quot; alt=&quot;alfred term&quot; /&gt;&lt;/p&gt;

&lt;p&gt;This fires up Terminal, which runs the chef search, and opens separate windows.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://farm6.staticflickr.com/5805/22422959375_891c2f313a_k.jpg&quot; alt=&quot;csshx cap&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Your cursor will be in the bottom red window by default and input will appear in all the windows at the same time. When done, CMD+Q will bring this window&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://farm1.staticflickr.com/744/22422971645_735d9473b9_b.jpg&quot; alt=&quot;quit window&quot; /&gt;&lt;/p&gt;

&lt;p&gt;And done :)&lt;/p&gt;

    &lt;p&gt;&lt;a href=&quot;https://blog.dbsmasher.com/2015/03/06/alfred-csshx-and-terminalception.html&quot;&gt;Alfred, csshx and terminalception&lt;/a&gt; was originally published by  at &lt;a href=&quot;https://blog.dbsmasher.com&quot;&gt;dbsmasher corner&lt;/a&gt; on March 06, 2015.&lt;/p&gt;
  </content>
</entry>

</feed>
