Using the scientific method to improve your AB testing

One advantage that Calm has had over its competitors is a strong culture around experimentation.   How did we build that and what goes into it?  The vast majority of articles on the role of data-science in AB testing discuss ways to make your testing process rigorous.  For example, you should run power analyses to determine how long to run your test before checking results for statistical significance.  What is rarely discussed, however, are methods for making your experimentational process as effective and efficient as possible. In other words, how do you figure out what the best version of your product is as quickly as possible?   In this article, I will discuss how to use the scientific method to achieve this and common pitfalls and misunderstandings that hold your experimentation practice back.  

Scientific method:  Research question, hypothesis, and prediction

To understand the scientific method, it is critical that you first properly understand the difference between a research question, a hypothesis, and a prediction.  The scientific process starts by determining what you want to learn about and posing it as a research question.  Why are so many of our users not converting?  Why are some people but not others forming a habit?  Generally “yes” or “no” questions should be avoided because they lead to very limited learnings – yes or no.  Try to formulate questions in a way so that if you answer it, the answer will have important implications.  Research questions sometimes arise from surprising data findings.  For example, we once saw that Calm users were meditating at the end of the evening, which isn’t an expected time for meditation, and wanted to understand why.  Often though, you already know what you want to learn about – something that you know will impact your business.

In the next step, you develop hypotheses, which are explanations or answers to your research question.  For example, people don’t know that they need to develop a habit to see the benefits of Calm or they want to develop a habit but keep forgetting to practice at established times. Hypotheses should be developed outside of the context of any specific test or prediction, so that you identify the best explanations, not just explanations that are easiest to test.  Good hypotheses should be testable however.  Good hypotheses also are not guesses. They should be guided by external research (e.g., cognitive research on habit formation), previous AB testing, marketing research, survey results, and data analysis. 

Once you identify a hypothesis you want to examine, you use it to make predictions.  Assume your hypothesis is correct, what would you expect to see and not expect to see?  If users are forgetting to meditate, if you enabled them to set up a reminder to meditate, you should see an increased number of them developing a meditation habit.  Tests are set up to see if you can confirm this prediction. The point of the test though is use that confirmation (or lack of one) to learn about the hypothesis. Good predictions are predictions that only your hypothesis would make, because, otherwise, if you confirm the prediction, you won’t have learned which one is correct.  Thus, you may be no closer to finding a good answer to your question.

Pitfalls and Misunderstandings

Not considering the wider possible set of research questions and hypotheses before prioritizing.  At the start, flesh out all the research questions you may want to answer.  If you can see all the possible options and prioritize them as a team, it will ensure that everyone agrees that you are trying to learn about an area with the most potential business impact and decreases the chance that you will waste time pursuing suboptimal questions.  Sometimes when you do this, it will help you realize that you should answer one question before addressing another one.

Once you identify the research question(s) you want to address, then spend time identifying a broad set of hypotheses.  This ensures that you do not simply prioritize ones that are the easiest to test or the first to come to mind.  Instead you want to prioritize hypotheses that are most likely to effectively address your research question.  Comparing different hypotheses also help you refine your hypotheses.  Does a person not convert because they cannot afford it or because they do not think it is worth as much as you charge?  If you consider only one of these hypotheses isolation, you may not realize that these hypotheses are distinct and lead to different predictions.  

Building out a list of potential hypotheses often helps you identify hypotheses that are conflicting.   For example, one hypothesis may be that users will find our product more valuable if it contains more content that can be passively consumed versus another that posits that users will find our product more valuable when it challenges them and makes them expend energy to grow.  These conflicting hypotheses lead to very different predictions and pathways in product development. Be on the lookout for contrasting ideas!

Failing to abandon a weak hypothesis.  Try to establish ahead of time what failure or set of failures would make you abandon a hypothesis.  It can be hard to be certain that a hypothesis is bad if one prediction is not confirmed – try as you may, tests are not perfect!  But perseverating on an idea that has little support decreases efficiency of your testing program.  In general, given opportunity cost, pruning down hypotheses, and moving on to new areas, is one of the most important learnings you can have as a company.   This is another good reason to have a prioritized list of hypotheses.  If you know that you have other good ideas to examine, you will have more motivation to abandon a weak hypothesis.

Not using data to evaluate and prioritize hypotheses.  Product tests, especially at start ups, can involve a couple months of work to move from the speccing stage all the way to final analysis and decision. One way to speed up your testing program though is to find alternative ways to assess the likelihood of a hypothesis before examining it.  You can run surveys, do market research or work with a user researcher.  You can also work with data analysts.  Make predictions about what you would expect to see in your current data if the hypothesis is true.  Analyses like these can be answered much more quickly than a test, sometimes in a couple hours, and will help you prioritize which hypotheses are likely and worth further pursuing.  

Outcome focused instead of learning focused.  Ultimately the goal of a testing program is to improve your critical business metrics, i.e, Key Performance Indicators (KPIs).  However, it is easy for researchers to become focused on KPIs and lose focus on learning about their users.  At worst, a company with a poor research culture will devolve into “throwing spaghetti at the wall to see what sticks,” running experiments with minimal rationale, with ideas based on simple optimization tweaks (e.g., change the font size!) to see if they move the metric.  

Teams that are focused on learning, however, are trying to understand their users – their needs, their motivations, and their problems.  They want to build knowledge up over time so that a series of experiments, together as a whole, effectively and efficiently resolve critical research questions.  They want to learn as much about what they should do as what they shouldn’t do.  They want to build up narratives about their users.  They are looking for deeper insights that can change the product direction.  For example, if you have a freemium product, instead of simply seeing if a product change successfully led to more conversions, you may want to figure out if users are most motivated to convert upon arrival or if they instead need to spend time using the free content in the app before they will convert.

Choosing a test metric based on desired outcome, not on what you want to learn.   If you are trying to obtain a specific outcome, such as positively impacting a KPI, you will use that KPI to assess whether your test was a win.  However, if you are trying to learn, you should use the metric that best enables you to assess whether your hypothesis is supported.  This metric may be the KPI but not necessarily.  For example, if you have a KPI that measures engagement but your test is specifically interested in habit formation, then use a metric that measures habit formation if your KPI doesn’t.

If you use a metric that isn’t your KPI, sometimes you will learn something important but not see a significant impact on your KPI.  This could happen for a number of reasons. It may need more statistical power to influence the KPI versus the metric you are interested in.  Sometimes the impact on the overall KPI is washed out when are are learning about a specific segment. Sometimes you learn a lever for influencing your customers but you need to pull harder on that lever before it has a major impact on your KPIs.

Treating predictions and hypotheses as synonymous.  If you don’t treat a prediction and hypothesis as distinct concepts, you will probably miss the deeper learnings enabled by a clearly defined hypothesis.  In turn, this means that it will be less clear how to iterate if your test is successful or not.  If you focus on hypotheses, your focus is on learning about your users.

Brainstorming product changes not hypotheses or predictions made by hypotheses.  Sometimes it can be helpful to brainstorm up a bunch of product changes that your users will like, for example, with a hackathon.  Although some fresh ideas can be helpful, this should not be your dominant approach to iterating on your product.  The problem is, again, that a product change or prediction that is not developed explicitly to test a hypothesis, will often result in you learning notably less.  Thus, over time, instead of your experiments building on each other, leading to new narratives and deep understandings, you will be more likely to run a scattered series of tests.  Ask yourself what you want to learn.  Develop your tests to learn it.  Avoid testing an interesting product change and then figuring out afterwards what you want to learn.

Developing tests that fail to address your hypothesis.  Experimental design is challenging.  It can be surprising sometimes that small changes to your variants or how you implement a test will dramatically change what you can learn from the test.  I’ve seen researchers waste a substantial amount of time running experiments that are not capable of helping them learn what they want to learn.  Consult with data scientists when designing a test.  If you make changes to the design after consulting with them, pass it back by them before implementing it. 

Final Recommendations

Brainstorm research questions and analytically prioritize them.  Then generate a set of hypotheses for the chosen research question.  Refine your hypotheses and search for competing ones.  Prioritize which ones work based on past learnings, theory, and/or data analysis.  Spend time at this step; it is impactful.  After you decide what you want to learn and which hypothesis you want to test, only then generate predictions and test ideas.   Tests should be designed so that you will learn about what you want to learn – your hypothesis.  The test metric should be the metric that is the best metric for evaluating your hypothesis.  Before running that test, think through, if the prediction is not confirmed, what other ways would you want to test your hypothesis before abandoning it.  Aim to fail if it isn’t supported.  If your hypothesis is supported, move on to other predictions and ideas suggested by your hypothesis.

M.I.C.E: Adding Motion to the I.C.E. prioritization framework

When I joined Calm in early 2017, we were drastically behind our main competitor in terms of key metrics like revenue, brand awareness and money raised.  In the next two years, with a team that was a tenth of the size of them, we managed to overtake them in most metrics and also to become the first mental-wellness unicorn.  We were amazingly still only a team of about 25 people a couple months before we signed term sheets on the valuation.   

How did we accomplish this?   Besides being a group of 10xers, one of our core strengths as a company has been an aggressive focus on prioritization.  “What will move the needle?”  We didn’t waste time on projects that wouldn’t drive growth and we were smart about the order in which we tackled the projects that would matter. 

On the data team at Calm, I’ve found it tremendously helpful to have a framework that makes it it possible to turn these prioritization questions into numbers and in a way that highlights critical factors impacting whether you should prioritize work.  In this article, I’ll discuss how we not only adapted the I.C.E. prioritization framework, but how we also made a useful modification to it – added the conception of Motion – and turned it into M.I.C.E. 

I.C.E.:  A prioritization framework

What is I.C.E.?  It stands for Impact, Confidence and Ease.  If you have two projects that will have an equal impact on the business’ bottom line but one is easier to complete than the other,  then do the one that is easy to complete.  You’ll get business value sooner and will have leftover time to work on another project.  This concept of working on the project that gives the most business impact per unit of time invested is also called leverage.  The ability to stay focused on the work with the highest leverage is a essential characteristic of a top-performer.

Confidence is an interesting component.  We can’t always be confident that the project will have the estimated amount of impact.  Multiplying impact by confidence is a way of assessing the expected value or the probable impact of a project.  If two projects are both deemed to have a high impact but you are only confident that one will have the intended outcome, then clearly choose the one with the higher expected value.  It can be helpful to address confidence and impact separately, in part, because sometimes, strategically, you want a balance of big bets – projects with high impact but high risk/low confidence – with projects that are more likely to work but have limited impact.  

To calculate the prioritization of projects, give a numeric score to I.C.E. (e.g., 1-4) and then calculate the average to obtain expected business value per unit of employee time.  Rank projects by their ICE score and then work on the projects with the highest ICE.  If you have projects with a similar ICE score, then, if the impact and confidence vary between the two, you could break the tie by determining if you want more big bets or clearer, medium wins.

Motion:  Adding an M to ICE

With my team at Calm, we realized that it was tremendously helpful to add Motion to ICE.  Motion captures how much energy there currently is behind a project on partner teams or the broader company.   Are other people ready to act on or use your work if you do it?  Do they know exactly what that action will be or do they only have vague ideas?  Does the project fit the company’s wider strategy and goals?  If you are assessing tech debt, does it fit with your team’s mission statement?   Does it improve infrastructure that will be needed for immediately subsequent projects? 

Motion matters, in part, because the vast majority of projects, particularly data projects, require a hand-off to partner teams before business value can be obtained.  Thus, if motion is low, even if the project does have plenty of potential, it is unlikely to lead to business value.  It doesn’t matter if your churn model could lead to high business impact, if partner teams like lifecycle or product teams are not interested or ready to use the predictions to address churn.  So if you don’t have motion, you may be making a mistake to prioritize that project. 

A useful characteristic of motion is that it is relatively easy, compared to the other factors, to change.  If it is low, but the ICE is high, then try to raise enthusiasm about the project and to get partner teams committed to working with you on the project.  If you can raise the motion, then reassess prioritization.  It may be at the top of your list now. 

Efforts to understand why a project has low motion can be enlightening and will usually point you towards action steps.  For example, if you find that there isn’t consensus around the project or that key people don’t understand the point of the project, then, if you believe in the ICE of the project, spend time pitching it to people.   If a partner hasn’t taken the time to think through what the concrete next steps will be when you hand off the work to them, then push them to do so.  If they know the project won’t get prioritized without those concrete plans, they’ll be motivated to develop them.   Sometimes your partner just needs time to clear other projects off their plate.  If that is the case, give it some time, motion will go up and so will the project’s prioritization score.  Note, sometimes you’ll realize that you are actually missing some context and that the motion is low for a good reason.  If so, reassess your scores for ICE and send that project to the bottom of your backlog!  

The conversations we have had with partner teams about motion have had good organizational consequences.  For example, it became clear when working with our UA team that some work done by data engineering was sitting for longer than expected, waiting for someone on UA to start using the new data.  We discussed how, in order to have good motion on these projects that we needed to improve the coordination in how we hand off the projects.  We subsequently took steps to ensure individuals in UA were prioritizing the same projects as data engineering in the same sprint.  

In sum, try scoring and ranking your projects with MICE.  You’ll stay more focused on higher leverage work and the conversations that arise around the discussion of motion, will help you spot problems and identify actions for improvement.  

** The idea for Motion was developed in collaboration with Ben Paul, a senior data scientist at Calm. Ben is excellent at identifying and focusing on high leverage work and turning analyses into impact.

SF Neighborhood recommender

When starting down the path of data after my post-doc, like a lot of other people, I worked on some personal projects to round out my skills and resume. One personal project, linked below, made it into an early edition of Data Science Weekly (www.datascienceweekly.org) and into the number one spot at datatau.net and r/datascience. The post also gain me a lot of attention, helping me land my first role as a data scientist. A few years later, the CEO of Mode used it as an example of an excellent data-science project (quora post). Even more important to me, I read a newspaper article years later that said that one of the SF neighborhoods noted as being undervalued in my blog post exhibited abnormally high growth of real-estate prices since then.

Why was this post successful and what can aspiring data-scientists learn from it? One major weakness that I often see in bootcamp projects is that they are purely technical. The aspiring data scientist cleans some data, applies an out-of-the-box algorithm it, and then maybe shows off the work in a web app or with a nice visualization. The problem is that these projects rarely demonstrate that you have enough technical strength to be a good hire based purely on your technical prowess (e.g., to be a machine-learning engineer). A boot-camp project isn’t really designed to demonstrate strength here either.

Data science projects that are successful, on the other hand, tend to demonstrate your ability to think through a vague problem and arrive at something useful. They show that you can use data-science methods – whether more technical ones like machine-learning models or simpler ones like data-munging and story telling – to elicit business value from data. They tell stories with data that make people want to act. They demonstrate creativity. They demonstrate that you will be thoughtful and rigorous in how you apply the data-science methods – that you don’t just copy and paste ideas or code from elsewhere, that you understand the nuances and trade-offs in your approaches, and that you are aware of business impact when you make decisions.

Best of luck to those working on breaking into the field!