That's a very weak blanket statement, there are totally reasonable A/B tests you can run that don't deteriorate a user's experience, and the results can guide you to a better customer experience overall.
It did not mean it too seriously, of course there are also good AB tests, but there are a lot of bad ones out there. Those are what the article was about.
Obviously, it depends on what the A/B test is about. Molesting or not the customer for the sake of some shortsighted metric is a bad choice; deciding what content should go above the fold or not (e.g. Amazon places images, short description, details/specs and similar products in that order) is a good choice.
AB testing can be (although isn't always) used to improve the customer experience. Assuming you know exactly what will make the customer experience best without actually testing it can also lead to a worse experience.
Maybe if you are very aware of the fact that your goal and metric aren't totally aligned, but that often gets lost. As a result A/B testing for longer website visits can make websites that make it more obvious that the information you want is there, but also make the path to actually get it longer. A/B testing for engagement might promote divisive behavior and fights. A/B testing for read rate or clicks might lead to trust loss.
I think a lot of lessons from AI safety apply surprisingly well to A/B testing, mainly around how hard it is to align your actual goals with the metrics you use for optimization, and how disastrous the consequences can be. It doesn't have to go wrong, but it's incredibly hard to ensure it goes right, especially if it's the only feedback you have.
I've spent a lot of my career doing A/B testing, including doing that role exclusively for a number of years. I specialize in ecommerce, so maybe I have too narrow of a view here, but in that vast majority of cases, I am optimizing for revenue per visitor, which is a function of conversion rate and average order value. There are sometimes leading indicators like engagement, but in ecommerce, you're afforded the luxury of basing things on revenue or even bottom line.
I really don't like the positioning of ALL A/B testing as unethical behavior where you're hostilely trying to take advantage of a user. It's quite the opposite. There are a lot of extremely poor user experiences out there and a quality testing program can help improve user experiences, remove risk from making sweeping changes, and help you learn more about your audience and market.
The vast majority of the successful testing I've done is done around trying to HELP users navigate the site and product catalog, understand the product, and purchase the product. Attention spans are fleeting with online shopping and even the smallest points of confusion or friction can turn shoppers off.
Additionally, often times I'll read into test results after a month or so to see if there were any issues with orders that might indicate purchases from disinterested people or misaligned expectations.
But that's just optimizing for bad metrics. At this point, anyone who thinks "engagement" and "time spent on page" are customer-positive metrics is in a different mental space than you and I. There's a lot of ineffable things that make up good customer experience that would be hard-to-impossible to A/B test, but it doesn't mean that A/B testing is "unsafe" just because it could be used to optimize for bad things any more than any other telemetry or metrics gathering could be bad because you could optimize for evil things. And at the same time, bad management and product leadership can optimize and develop towards bad goals with plenty of tools that aren't A/B testing.
It seems to miss the point to blame/stigmatize a specific tool because it's been used poorly by a few bad actors in a public way.
Further, I imagine that the obvious "known bad" metrics are not selected only by "A few bad apples". I think it's likely they are selected by the mass of business actors looking for current quarter results.
For sure. I don't think there's general "overall" metrics that you want to be testing against every single time on every change outside of basic performance metrics for loading or rendering in real-world environments.
I wasn't at all trying to say that only a few places are optimizing for bad things, but as you see all over this thread, there's a number of companies that immediately come to people's minds as bad actors when it comes to A/B testing - Google, Meta, Microsoft. There's plenty of other companies that are more ethical about it, or use it as part of rolling out general changes and collecting feedback. I know half of the time I log into the AWS Console it has some sort of "Hey, we're testing out a new upcoming UX for this page. Click here if you want to go back to the old one", which seems like a decent way for them to get feedback on the new designs while not drastically disrupting things.
Bad management can certainly ruin things without A/B testing.
It doesn't excuse A/B testing simply being a poor tool among all you have access to. Talking to users and stakeholders, for example, provides infinitely more input. (Edit: yeah in many cases measuring what users do, directly watching or via analytics, is also useful.)
Definitely - I'm not trying to say A/B testing is amazing, just that a lot of the comments have a strong "if you do A/B testing you're evil and are out to manipulate people" bent to them, which I think is too far in the other direction.
Talking to people is great, but getting a representative sample is hard, and often people are bad at both understanding what they want, expressing it, or even being accurate about how they use things. I know when I was working closer to the UX side of the business before, I was constantly surprised by both what users would say they want AND by how users actually used the products.
In my mind, A/B testing is good as a sort of "final pass" to serve as broad, semi-random validation that the change you're looking to make does actually do the thing that it's intended to do. It's not great for early on when you don't really know what to measure or look for, or if the change is remotely reasonable, but it can help check for if your focus group/user panel happened to be weirdly skewed in their usage/desires.
They're only different if you've selected bad metrics. If you've got two different search algorithms, running an A/B test and measuring how often the user selects the first item returned is a good measure for how well your search algorithm is returning the information the customer wanted, which is good customer experience.
They are always different. You cannot hold a conclusive A/B test for customer experience.
Search engine, a single-purpose tool, is as simple as they come regarding customer experience. Still, a good search algorithm can make me click on the first result if it is good, and a bad search algorithm can make me click the first result because they are so bad that scrolling further is a waste of time, especially if I already needed to scroll through widgets and ads to get to the first result.
It's not about just selecting good metrics, it's about higher level picture that A/B testing can never get you.
> Assuming you know exactly what will make the customer experience best without actually testing it can also lead to a worse experience.
For that you usually hire a market research company or do what they will do: take an interviewer, two cameras (one front-face, one top-hands) and hire an as-diverse-as-possible pool of test candidates that you then put through whatever workflow optimization you want to do. Then afterwards, you interview them - side benefit, you can get really interesting general side knowledge that you'd never gain from a dumbass A/B scheme: is your font style/color scheme legible, can the site be used by colorblind people, are there stock photo choices that give off stereotypical vibes...
It's real fun and a worthwhile experience for everyone involved.
It's not really an either/or option. You can use testing to validate the changes stemming from market research.
Having seen lots of site redesigns go horribly wrong due to 100% earnest people trying their best and utilizing the research that was afforded to the process, I always recommending incrementally testing into changes on high-value / high-risk applications, even when the "improvements" were backed by solid research. You never know until you release.
The "or" was meant to be the distinction on who does the user testing - I've seen both in-house testing operations and outsourced ones. For small scale operations, it may actually be cheaper to run them in-house and only hire external testers... cameras are dirt cheap these days.
Hiring a market research company is usually worth it if you have a contract with them anyway (which gives you better rates on the testing) or lack someone on staff who knows how to deal with cameras.
I have experience where the company paid a UX agency to create a flow that was by all standards better customer experience and a better product, nicer too. They ran an AB test, turns out people were more likely to pay with the old version. AB testing is good that it challenges what UX people think is better experience or product with hard metrics.
Which is why A/B testing is an important part of the UX toolkit. It's a tool among others, and is one way to validate assumptions. A good UX designer will try to base their designs on data and reasonable hypotheses drawn from the data, but a new design or flow is necessarily based on some amount of assumptions, so it requires validation.
That said, an A/B test does not tell you why something didn't work. You can make further assumptions based on the results and develop new hypotheses, but it never tells you why. Typically you would do some kind of qualitative UX research on a prototype or even static concepts beforehand to identify these kinds of issues before you even expend the effort to do a live A/B test. Far cheaper to do a study with 6-12 people and a prototype than to build out a full, functioning A/B test experience.
It's possible the flow they created was generally better but perhaps it had one fatal flaw. Perhaps that flaw could easily be remedied once identified.
A/B testing is just one small part of a good UX process.
There are many other metrics that do not involve AB testing. You can just survey customer experience before, during and after a purchase for example. I never said to throw out all metrics.
With AB testing your are optimising for a specific outcome. Usually higher conversion. As pointed out in the article eventually you'll end up with a bunch of colourful buttons and scary texts that persuade the user to click. A lot of the "only 2 seats/rooms available" are lies to scare the user into a conversion.
AB testing is how you isolate a change and measure the impact. It's the only real way to be able to associate cause and effect. Best you can do otherwise is measure something over time while making changes. You can try and correlate changes with outcomes but it's hard to be sure the change is what drove the outcome.
That sounds pretty accurate. Anyone who's gone through some econometric classes would know that split testing is (almost) the only way to get to (close to) causality - all the rest is black magic called correlation which can also lead you to the conclusion that Nicolas Cage's movies lead to people drowning:
Instead try to improve the customer experience, make better products, improve customer service.