Search Russell Research

Find a solution, case, or article

Enter at least two characters to search the site.

    Press Esc to close. Use Ctrl/⌘ K to search from anywhere.

    Advanced Analytics

    Data Science & Integrated Analytics

    A customer may say one thing in a survey and do another in the account record. Russell's researchers and statisticians connect the relevant sources, choose the analysis, test its assumptions, and say where the conclusion is strong or uncertain.

    Research design

    Start with what the data can answer

    Survey responses can explain what customers say matters. Account records can show purchases, renewals, or service use. Before combining them, we check whether they describe the same people, behavior, and period.

    We begin with the decision and an inventory of the available sources. Definitions, coverage, and missing information determine whether the work can explain a relationship, estimate a future outcome, or compare scenarios. Sometimes the useful next step is a narrower question or additional research.

    Define the outcome
    Agree what is being explained or estimated. An intention to renew and a recorded renewal answer different questions.
    Check coverage and timing
    Identify who appears in each source, who is missing, and whether the measures cover comparable periods.
    Establish a valid connection
    Review identifiers, shared variables, and permitted uses before deciding whether sources can be linked or should be compared separately.

    A relationship in the data does not by itself establish cause and effect. A before-and-after comparison also needs to examine what else changed.

    What Does Combining Data Mean?

    The connection determines which questions the analysis can answer. These three approaches provide different kinds of evidence.

    Link the same records

    Can we connect a response to a recorded outcome?

    A reliable shared identifier can connect survey answers with the same customer's records, where the agreed uses permit it. Unmatched records and differences in coverage still need checking.

    Compare groups or periods

    Do the sources describe comparable audiences?

    Without a shared identifier, sources may be compared by compatible groups, places, or periods. This can show differences at that level; it does not connect an individual's answer to their behavior.

    Estimate a connection

    Do shared variables support statistical matching?

    When the available measures and assumptions support it, statistical matching can estimate relationships across sources. Those estimates need validation and must remain distinguishable from observed records.

    Illustrative question: Do customers who say they intend to renew later renew their account? Answering it for the same people requires a valid link and a later outcome. Separate survey and account totals alone cannot establish that relationship.

    Matching the analysis to the question

    Choose the method against the outcome, available measures, and intended use. An analysis that explains existing differences may need a different design from one used to predict future behavior.

    Two people reviewing printed charts beside a laptop displaying a bar chart

    Method

    Driver analysis

    When it helps

    Examine which measured factors are associated with an outcome, accounting for the other variables included. Associations help identify explanations to investigate; they do not establish causes.

    Business decision

    Which factors are most closely related to satisfaction, consideration, or another defined outcome?

    Method

    Predictive modeling

    When it helps

    Estimate an outcome using information available at the point of prediction. Evaluate performance on separate data that reflects the intended use.

    Business decision

    How well can the available information estimate a future outcome?

    Method

    Advanced segmentation

    When it helps

    Identify groups using relevant needs, attitudes, or behaviors. Check whether the groups are distinct, interpretable, and usable with the information available to your team.

    Business decision

    Which audience differences warrant a different offer or approach?

    Method

    Data fusion

    When it helps

    Assess how separate sources can be connected, compared, or statistically matched. Document coverage, connection rules, and estimated relationships.

    Business decision

    What can we learn together that each source leaves unanswered?

    Method

    Simulation

    When it helps

    Compare defined options using a model and explicit assumptions. Examine how the comparison changes when important assumptions vary.

    Business decision

    Which option performs better under the conditions being modeled?

    Method

    Text analytics

    When it helps

    Organize open-ended feedback into themes, with researcher review of coding and examples. Interpret frequency against the people and comments represented.

    Business decision

    Which themes recur, and how do they differ across the feedback included?

    What Your Team Can Use

    Agree the outputs around who will use the analysis and whether it supports a single decision or repeated updates. The explanation and limits belong alongside the result.

    Integrated analytical dataset For research and analytical teams reviewing the evidence

    An analysis file with documentation of the sources, definitions, and connection rules used. Its structure depends on whether records are linked directly, compared in groups, or statistically matched.

    • Source coverage, observation periods, field definitions, and agreed uses.
    • Matching rules, unmatched records, missing information, and relevant exclusions.
    • A clear distinction between observed measures and estimated values.

    Illustrative question

    Which customers are represented in both sources, and how do those left out differ?

    Findings, model, and validation For decision makers and colleagues reviewing the analysis

    A plain-language explanation of the result, supported by the model diagnostics and technical detail needed to assess it. The validation should reflect what the model will be used for.

    • Estimated relationships and an explanation of what they do and do not establish.
    • Relevant performance checks, errors, and differences across groups.
    • A technical appendix documenting methods, assumptions, and sensitivity checks.

    Illustrative question

    Does the finding hold when the period, audience definition, or analytical choices change?

    Scenario comparisons For teams evaluating an offer, price, or other proposed change

    Tables, views, or an agreed simulator compare the options under stated assumptions. Modeled outcomes should remain distinguishable from observed results.

    • The options, inputs, and conditions used in each comparison.
    • Estimated outcomes and how they respond to changes in key assumptions.
    • Guidance on which comparisons the model supports and what lies outside its scope.

    Illustrative question

    Does the preferred option change when a key assumption is less favorable?

    Scoring and typing tools For teams planning to apply or update the analysis

    Where included in the scope, a scoring file or classification routine applies the model to suitable records. Repeated use requires agreed inputs, ownership, and checks.

    • Required fields, scoring or classification rules, and output definitions.
    • Validation appropriate to the intended audience and application.
    • Instructions for updates and checks when data or behavior changes.

    Illustrative question

    Can the team apply the model with the information it will actually have at the time?

    Data and Modeling Questions

    Can Russell analyze data we already have without commissioning a new survey?

    Yes, when the existing data can address the question. We begin with the decision you need to make, the outcome to understand, and an inventory of available sources. Survey files, customer records, transactions, engagement data, and open-ended feedback may each contribute different evidence. We review how they were collected, whom they cover, and what the fields mean before recommending an analysis. New research may still be needed if the available information omits a relevant audience, motivation, or part of the decision.

    How do we know whether our data is sufficient for the analysis?

    The number of records is only part of the assessment. We review missing values, duplicates, inconsistent definitions, time coverage, and whether the outcome and relevant groups are represented adequately. For a predictive model, the data must also contain information that would be available when a future prediction is made. Missing information may require a narrower question, additional data, or an explicitly qualified estimate. Filling gaps or adding model complexity does not resolve a mismatch between the data and the decision.

    Can datasets be combined if they do not share a customer identifier?

    Sometimes, but the type of connection determines what can be concluded. A reliable shared identifier may support record-level linkage when the permitted uses allow it. Without one, sources may still be compared by a compatible period, geography, or audience group. Statistical matching may be considered when common variables and assumptions support it, but estimated relationships must remain distinguishable from observed records. We assess the connection before treating separate sources as evidence about the same person or event.

    What happens when survey findings and customer records disagree?

    First, check whether they describe the same audience, behavior, and period. A survey may ask about category purchases while an account record covers only purchases from your business; an intention and a completed action are also different measures. We examine wording, coverage, recall, timing, and any record-linkage limitations before interpreting the discrepancy. Some differences reveal useful questions about customer behavior. Others reflect incompatible definitions or missing coverage. The analysis should explain the difference rather than force the sources to agree.

    Can driver analysis tell us what will cause an outcome to improve?

    Driver analysis can identify factors associated with an outcome and help prioritize explanations to investigate. Those relationships do not by themselves establish what would happen if the business changed a factor. Other influences, differences between groups, or the timing of measurement may explain part of the relationship. If the decision requires evidence of an effect, we assess what experimental or comparison design is feasible. A before-and-after change or a strong predictor alone is not sufficient to establish causation.

    How do you check whether a model will work beyond the original data?

    Validation should reflect the model's intended use. For predictive work, that includes evaluating performance on data that did not guide model fitting or selection, with separation by time or customer where the application requires it. We also examine errors, performance across relevant groups, and sensitivity to analytical choices. A good fit to the original records is not enough. The readout should explain where the model performs adequately, where it struggles, and what still needs checking before it informs ongoing decisions.

    How will our team understand the model and its limits?

    The explanation should connect the business question to the data, method, assumptions, and result. We provide the diagnostics and sensitivity checks appropriate to the analysis, alongside a plain-language account of what the evidence supports. Technical detail can sit in an appendix while the main readout focuses on the decision and its uncertainties. We distinguish stable findings from conclusions that depend on a particular definition, model choice, or scenario, so your team can judge how much reliance to place on each recommendation.

    Can our team continue using the analysis after the project?

    That should be planned during scoping. A one-time decision may need documented scenario tables and a readout; repeated use may call for a scoring file, classification routine, simulator, or dashboard. We agree what inputs the output requires, who will maintain it, and what checks should accompany an update. Changing source definitions, customer behavior, or market conditions can affect its usefulness. Producing a model or tool does not by itself establish a maintained connection to your operational systems; that implementation needs its own agreed scope.

    Plan your research

    Explore tools selected for this research. Browse all research tools and calculators.

    Discuss the Data and the Decision

    Tell us what sources are available and what you need to understand. We will assess what can be joined, what analysis is appropriate, and what new data may still be needed.

    Talk with a Researcher