Aydınlatma metni yükleniyor…
A/B testing is not simply a matter of comparing two versions of a web page and selecting the one with the higher conversion rate. A reliable experiment must connect the business objective, user behaviour, variation assignment, analytics events, technical performance and decision rules within one coherent architecture. Without that foundation, it can be impossible to determine whether a change in results came from a new headline, button or form—or from a measurement error, traffic shift or implementation defect.
Conversion rate optimisation, or CRO, aims to improve the rate and quality of measurable actions such as form submissions, quotation requests, product views, account registrations and purchases. For marketing, digital product, e-commerce and technology teams, the objective should not be to test isolated ideas at random. It should be to create a repeatable experimentation system that identifies friction in the user journey, measures commercial value and produces knowledge that can guide future decisions.
What is an A/B testing architecture?
An A/B testing architecture is the technical and operational framework that assigns eligible visitors to a control or variation, delivers a consistent experience to each group, measures behaviour using the correct identity and evaluates results according to predefined rules. The control represents the current experience, while the variation introduces a specific change intended to test a hypothesis.
A dependable architecture has five connected layers: objectives and hypotheses, user and segment definitions, variation assignment, measurement and data, and the process for reaching a decision and releasing the result. Design, development, analytics and marketing teams must work from the same definitions. Otherwise, a form may display a success message without recording a valid submission, or the same person may encounter different variations across separate visits.
Start with a business objective, not a design change
The first question should not be “What should we change?” It should be “Which business outcome and user problem are we trying to improve?” A completed quotation form on a B2B website, an add-to-basket event on an e-commerce platform and a new account in a subscription product do not carry the same decision value. The experiment must connect to a meaningful outcome such as revenue, qualified demand, transaction completion or user activation.
A useful hypothesis clearly states the observation, proposed change, target audience and expected effect. For example: “If we move delivery information closer to the purchase button on mobile product pages, the add-to-basket rate among new visitors will increase because there will be less uncertainty at the point of decision.” This statement limits the design scope, audience and primary metric. If the result is negative, the team can also identify which assumption was not supported.
Experiment candidates should come from evidence rather than brainstorming alone. Web analytics, funnel performance, site-search records, form-error logs, customer-service feedback, user research and technical performance data can reveal different parts of the same problem. GOV.UK’s guidance on usability benchmarking highlights the value of interpreting performance measures alongside user research, particularly across end-to-end service journeys.
How to prepare the measurement plan
Choose one primary metric
The primary metric determines whether the experiment has achieved its stated objective. It might be the rate of successful form submissions, completed purchases or qualified quotation requests. A button click should not become the primary metric merely because it is easy to capture. If users click but fail to complete the subsequent process, the apparent improvement may have little commercial value.
The denominator must also be explicit. A conversion rate calculated across all visitors will differ from one based on eligible users, sessions or unique users. The measurement plan should define the unit of analysis and explain how repeated visits, devices and authenticated accounts are treated.
Add secondary and guardrail metrics
Secondary metrics help explain why a change worked or failed. Examples include form starts, field-level errors, progression to product details and average order value. Guardrail metrics protect outcomes that should not deteriorate in pursuit of conversion. Cancellation and return rates, low-quality leads, page performance, application errors and accessibility problems may all belong in this category.
A variation that increases form submissions but substantially reduces lead quality should not automatically be considered successful. Where the journey continues into a sales operation, a well-planned website–CRM integration makes it possible to evaluate sales acceptance, follow-up and opportunity quality instead of treating every submission as an equal conversion.
Before launch, document event names, parameters, trigger conditions, ownership and validation rules in a measurement dictionary. A form-submission event should normally be recorded only after the server confirms success. Retries and double clicks must be deduplicated, and the experiment ID, variation ID and assignment time should travel with the event wherever practical.
User assignment and meaningful segmentation
Eligible visitors should be assigned randomly to the control or variation. The assignment key may be an authenticated user ID, an anonymous first-party identifier or another persistent identity suited to the product. If the same person cannot be recognised across devices, the likely analytical impact should be documented rather than ignored.
Assignment should happen before the experience is displayed and remain stable throughout the test whenever possible. Users switching between versions can dilute the measured effect and create an inconsistent experience. Authentication, consent changes, cookie deletion and cross-domain journeys therefore need explicit handling rules.
Segments should relate directly to the hypothesis. New and returning visitors, mobile and desktop users, traffic sources, customer types or countries may represent meaningful differences. However, searching dozens of segments after seeing the overall result increases the chance of finding an accidental winner. Predefined segments should inform the main analysis; patterns discovered afterwards should become hypotheses for future experiments.
An unusual difference between the planned traffic split and the observed allocation can indicate an implementation or data problem. Bot traffic, employee visits, analytics blockers, consent status and redirects may also affect eligibility. Include clear inclusion and exclusion criteria in the experiment record, then monitor the allocation before relying on the outcome.
Client-side or server-side testing?
In a client-side implementation, code running after the page loads transforms the control interface into the variation. This approach can be quick to deploy, but it may briefly expose the original content, cause layout shifts, add JavaScript weight or create analytics timing problems. These effects are especially important when the experiment changes prominent content above the fold.
With server-side testing, the variation is selected before the response reaches the user. It generally offers a more controlled basis for experiments involving business logic, pricing, search ranking, personalisation or checkout. It can also keep assignment closer to authenticated identity and server-side transaction records.
Google’s field Web Vitals measurement guidance recommends connecting experiment groups to analytics data and notes the advantages of determining groups on the server for performance analysis. This does not make one technology mandatory for every website. The appropriate choice depends on the experiment’s scope, content management system, caching model, team capabilities and acceptable performance cost.
In either approach, the data layer should expose the experiment ID, variation ID and assignment timestamp. CDN caching, session management and personalisation rules must prevent one variation from leaking into another. The testing layer should be planned as part of the website’s product architecture, including its management interface, analytics, CRM and e-commerce services—not as an isolated script added after launch.
| Stage | Control | Risk | Evidence |
|---|---|---|---|
| Planning | Primary metric and minimum detectable effect | Wrong objective | Approved hypothesis |
| Implementation | Persistent random assignment | Mixed experiment groups | Assignment record |
| Measurement | Event and conversion validation | Missing or duplicate data | QA report |
| Release | Performance and accessibility | Experience degraded for conversion | Guardrail metrics |
| Decision | Effect size and confidence interval | Premature or incorrect decision | Experiment result record |
Sample size, test duration and statistical decisions
Before the experiment begins, estimate the required sample using the baseline conversion rate, minimum detectable effect, accepted error risk and expected eligible traffic. Detecting a very small improvement requires more traffic and usually a longer test. On low-traffic pages, micro-conversions may provide useful diagnostic evidence, but they should not be presented as substitutes for the final business outcome.
Stopping after a few strong days increases the risk of a false conclusion. Campaign launches, weekday and weekend behaviour, salary cycles, stock changes and B2B purchasing schedules can influence results. Duration should therefore reflect relevant business cycles as well as the sample target. Teams should also avoid repeatedly checking the data and stopping as soon as a preferred result appears unless their statistical method explicitly supports sequential decisions.
Statistical significance is not the same as commercial significance. Microsoft’s experimentation guidance explains how p-values relate observed evidence to the null hypothesis and why noise can lead to misleading interpretations. Instead of reporting only that a variation “won,” present the effect size, confidence interval, sample, duration and guardrail results. A statistically detectable change may still be too small, risky or expensive to justify permanent implementation.
Quality assurance, privacy and data security
Before launch, inspect both experiences across relevant screen sizes, browsers, authentication states and traffic sources. Test form validation, payments, redirects, analytics events, error messages, keyboard navigation and screen-reader behaviour. A controlled rollout to a limited share of traffic can expose serious defects before they affect the entire audience.
Quality assurance should also verify that users remain in the same variation, events carry the correct experiment attributes and conversions are not duplicated. Compare browser events with server or transaction records where possible. If those sources disagree, investigate the discrepancy before interpreting uplift.
The experimentation system should not collect unnecessary personal data. Define the purpose, data-minimisation rules, retention period, access permissions and deletion process. If cookie consent affects measurement, document how non-consenting visitors are treated in both the technical design and analytical method. Testing platforms, tag managers and third-party integrations should be included in security and performance reviews.
Move the result into the permanent product
Leaving a winning variation active indefinitely inside a testing platform creates technical debt. Once approved, the change should be implemented in the core codebase or content management system, the temporary experiment code should be removed, and post-release metrics should be monitored. This confirms that the permanent implementation reproduces the observed effect without introducing regressions.
Negative and inconclusive experiments are also useful. They can reveal an unsupported assumption, an effect too small to matter or a measurement weakness. Hiding these results encourages teams to repeat failed ideas and creates an exaggerated picture of the programme’s success.
A central experiment register should contain the hypothesis, owner, variations, audience, metrics, start and end dates, technical version, result and final decision. This preserves knowledge when teams change and enables new tests to build on prior evidence.
Priorities for a sustainable CRO programme
- Score experiment candidates by expected business value, user impact, strength of evidence and implementation cost.
- Identify tests that could overlap within the same journey and establish mutual-exclusion rules.
- Approve primary, secondary and guardrail metrics before the experiment begins.
- Validate analytics events against browser, server and transaction records where appropriate.
- Store winning, losing and inconclusive experiments in the same shared knowledge base.
- Monitor performance, accessibility, security and data quality as consistently as conversion.
- Define how approved changes will move from the testing platform into the maintainable product.
Kumsal Agency is an Istanbul-based digital agency bringing together brand-specific web design, custom web development, information architecture, user experience, digital brand consulting and e-commerce. Its approach combines creative experiences that reflect the brand with business goals, integration requirements, performance, manageability, data security and modern software architecture.
For corporate websites and e-commerce platforms, experimentation can be designed as a sustainable part of the product architecture. The objective is not merely to prepare a different interface. It is to connect creative design with trustworthy measurement, manageable technology and real commercial outcomes.
Build an A/B testing roadmap for your website
Successful CRO depends less on producing more variations than on building a system capable of answering the right questions reliably. Assess your website’s conversion goals, existing data infrastructure and priority user journeys together. Contact Kumsal Agency to create a project-specific A/B testing roadmap aligned with your brand, technology and business objectives.


