Skip to content
TablePage.ai Open the app

How to Build a Defensible Location Comparison From Public Data

Build trade areas, align ACS and CBP data, score candidates transparently, stress-test assumptions, and inspect finalists before committing.

Share X in f
Wei Hu

A defensible site selection analysis begins with business requirements, not a demographic download. Define the customer and operating model, set minimum requirements, screen markets, construct realistic trade areas, collect comparable demographic and economic data, score candidates transparently, test alternative assumptions, inspect finalists, and then approve, reject, or request more evidence.

The goal is not to find the highest unexplained score. It is to identify candidates that remain viable after uncertainty, costs, operational constraints, and field conditions are considered. A high score is a decision aid, not a guarantee of profitability or site success.

Start with the decision, not the available data

Site selection analysis is a structured comparison of candidate locations against measurable business requirements. Those locations may be regions, trade areas, or individual properties, but they should not be treated as interchangeable units of analysis.

Separate the work into three levels:

  1. Regional market screening identifies broadly suitable states, metropolitan areas, counties, or cities.
  2. Trade-area analysis evaluates the customers, workers, competitors, and access conditions around a proposed location.
  3. Property-level due diligence determines whether a particular parcel or building can satisfy operational, legal, physical, and financial requirements.

A county may match the target demographic while a proposed building has poor access, inadequate utilities, prohibitive buildout costs, or incompatible zoning.

Use this workflow:

  1. Define the target customer, operating model, and desired outcome.
  2. Translate those requirements into measurable criteria and minimum thresholds.
  3. Screen regional markets.
  4. Define a suitable trade area around each candidate.
  5. Collect aligned demographic, economic, competitive, accessibility, property, and operating data.
  6. Normalize and score the candidates.
  7. Test different boundaries, weights, vintages, and uncertain values.
  8. Inspect finalists using a standardized field checklist.
  9. Approve, reject, negotiate, or request additional evidence.

The criteria and weights must reflect the business model. A retailer may emphasize customer origins, visibility, parking, nearby destinations, and pedestrian or vehicle activity. A service business may care more about appointment demand and technician travel time. An office may prioritize workforce access, transit, and occupancy cost. A logistics facility may emphasize highway access, delivery windows, utilities, labor, and site configuration. One universal variable set would conceal these differences.

Create decision gates as well as scores. A candidate might advance from regional screening only if it passes a minimum market-size requirement, from desktop review only if access and property constraints appear workable, and from field review only if expected costs fit the break-even model. A composite score should inform these gates, not trigger automatic approval.

Build a source-to-variable plan for market research

Before gathering data, create a source-to-variable registry. It should contain eight fields: decision criterion, proposed variable, source, geography, reference year, coverage limitation, uncertainty field, and refresh date. Splitting the registry into two linked tables keeps it readable without losing those controls.

Decision criterion Proposed variable Source Geography
Customer base Target-age households ACS Trade area or component Census areas
Workforce Education or commute measure ACS County, tract, or block group
Industry base Establishments by NAICS industry CBP County or ZIP-based geography
Employment scale Industry employment CBP County
Competition Active competitor locations Verified business or POI source Candidate trade area
Access Travel time, traffic, transit, or parking Transport or mobility source Network or site
Property feasibility Rent, area, utilities, and zoning Broker, owner, utility, and authority records Property

Use the same row order in the provenance section:

Reference year Coverage limitation Uncertainty field Refresh date
Selected ACS vintage Estimate; custom boundaries require a documented transformation Margin of error Scheduled annual review date
Selected ACS vintage Precision varies by product and geography Margin of error Scheduled annual review date
Selected CBP vintage Employer establishments only Suppression or status field Scheduled annual review date
Selected CBP vintage Reference-period employment measure Suppression or status field Scheduled annual review date
Observation date Closures and classifications may lag Verification status Monthly or quarterly review date
Observation period Method, sample, and coverage vary Coverage or confidence field Monthly or quarterly review date
Quote or decision date Terms may be conditional Verification status Before each decision gate

Replace generic schedules with actual dates in the working registry. A field such as “review annually” does not show whether a review is overdue, while a recorded next-review date does.

The American Community Survey is the main public source for local demographic, social, economic, housing, education, employment, language, commuting, and workforce estimates. Its values are estimates rather than exact counts, so retain the estimate, margin of error, geography, product, and vintage in the working dataset. The Census Bureau’s ACS resources for businesses describe relevant subjects and uses for evaluating target markets and workforce characteristics.

Margins of error matter when comparing small areas or candidates with similar estimates. Do not report a narrow apparent lead as decisive when the available uncertainty fields show that the ordering may be unstable. Preserve the original uncertainty data so reviewers can distinguish measured differences from false precision.

County Business Patterns complements the ACS with annual establishment, industry, employment, and payroll measures. Its coverage is limited to establishments with paid employees, and employment refers to the week of March 12 rather than a full-year average, according to the Census Bureau’s County Business Patterns program description. It should not be interpreted as a complete count of all businesses or self-employed activity.

Census data will not resolve every decision. Source current competitor locations, traffic or mobility conditions, rents, buildout costs, zoning, taxes, utilities, infrastructure capacity, incentives, and operational restrictions separately. Before joining sources, verify that their vintages, geographic definitions, reference periods, and industry classifications are compatible. Where they are not, record the transformation and its limitations instead of hiding it inside a spreadsheet formula.

Define the trade area before comparing demographics

The selected boundary determines which people, businesses, and competitors enter the analysis. Choose and document it before comparing demographic totals.

Boundary method Best use Main advantage Main limitation
Fixed-radius buffer Early screening Simple and consistent Ignores networks and barriers
Drive-distance polygon Distance-sensitive services Follows reachable roads Equal distance may require unequal time
Drive-time polygon Convenience and service access Reflects network travel time Depends on traffic and routing assumptions
Customer-origin catchment Analysis using comparable operations Reflects observed behavior Requires representative operating data

A fixed radius is a useful preliminary screen, not proof of realistic market reach. Roads, transit, rivers, railways, street connectivity, parking, competing alternatives, delivery limits, and customer willingness to travel can all change a catchment.

Drive-distance and drive-time polygons use the transportation network. Document the routing and traffic assumptions. Customer-origin catchments instead depend on observed customer records; check whether those records represent the proposed operating format and market.

GIS layers can combine customer or sales locations, ACS demographics, competitors, complementary points of interest, accessibility, and location performance. This makes barriers, clusters, coverage gaps, and overlap easier to inspect. Heat maps and demographic overlays still show patterns or associations; they do not establish demand, causation, or future performance by themselves.

If a custom trade area cuts across Census boundaries, document the allocation or transformation method, the source geographies, and the assumptions introduced. Test whether a plausible alternative transformation changes the comparison. Do not present an allocated estimate as though it were a direct observation for the custom boundary.

Create a transparent candidate-site scorecard

The summary scorecard should keep each decision dimension visible rather than burying several concepts in one cell. Because a wide table becomes difficult to audit, use linked component tables with one row per candidate.

Candidate Market demand Demographic fit Workforce
Site A Score and key evidence Score and key evidence Score and key evidence
Site B Score and key evidence Score and key evidence Score and key evidence
Candidate Competition Access Infrastructure
Site A Score and interpretation Score and constraints Score and constraints
Site B Score and interpretation Score and constraints Score and constraints
Candidate Property feasibility Cost Unresolved risks
Site A Pass, conditional, or fail Score and assumptions Gaps and blockers
Site B Pass, conditional, or fail Score and assumptions Gaps and blockers

Also display the major decision subtotals separately:

Candidate Market attractiveness Site quality Financial feasibility
Site A Subtotal Subtotal Subtotal
Site B Subtotal Subtotal Subtotal

Keep a long-form metric table behind this summary. For every candidate and metric, record:

  • raw value;
  • source and source URL;
  • geography and trade-area method;
  • vintage or observation period;
  • margin of error or another uncertainty field;
  • preferred direction;
  • normalized value;
  • weight;
  • weighted result;
  • analyst note.

Choose a normalization rule before combining metrics. For example, map each criterion to a documented 0–100 scale using fixed benchmarks. Document the selected method, including the treatment of missing values, outliers, and ties. Reverse the direction of cost or risk measures where necessary so that higher normalized values have a consistent meaning.

Weights should map to documented business requirements, not be described as universal or empirically proven. Check related measures for double counting. Population, households, density, visits, income, and spending capacity can capture overlapping concepts; assigning each a full independent weight may allow one underlying idea to dominate the model.

Use explicit benchmarks, such as a minimum viability threshold, a relevant local market average, or comparable existing locations. Record why the benchmark fits the decision and whether a candidate must meet it or is merely being compared with it.

Run at least three scenarios: demographic-heavy, accessibility-heavy, and cost-heavy. Before calculating them, define what constitutes a material change. The rule could flag a candidate that crosses an approval threshold or moves enough positions to alter which sites advance. The appropriate rule is decision-specific, but it should be documented before the results are known. A stable second-place candidate may be more defensible than a nominal winner that falls sharply under plausible assumptions.

Test competition, whitespace, and cannibalization claims

Competitor counts need an interpretation tied to the business model. Do not assign them an automatically negative value.

Test the interpretation against population, category behavior, accessibility, customer movement, and performance at genuinely comparable locations.

Map existing locations, competitors, and their catchments to identify coverage patterns and possible geographic conflicts. Trade-area overlap is an initial cannibalization warning, not a standalone estimate of transferred revenue. If the business already operates nearby, compare customer origins, sales patterns, visit timing, and overlapping catchments before treating the proposed site as entirely incremental demand.

Tutorial datasets can demonstrate the mechanics of points-of-interest mapping, sales filters, and regional comparisons. For example, Esri’s regional-market GIS tutorial explicitly identifies its branded scenario as fictitious. Conclusions from tutorials and vendor datasets should remain illustrative until validated against the organization’s own operating evidence.

Audit uncertainty, bias, and real-world feasibility

Use a documented quality checklist for every material input:

  • Recency: Is it current enough for the decision?
  • Geographic fit: Does the boundary represent the relevant market?
  • Coverage: Who or what is excluded?
  • Reference period: Is it a point-in-time, weekly, annual, or multi-year measure?
  • Consistency: Are definitions stable across candidates?
  • Uncertainty: Are margins of error, suppression, and confidence fields retained?
  • Representativeness: Does the observed sample match the target population?
  • Licensing: Is the intended use and republication permitted?
  • Privacy: Is the data sufficiently aggregated?
  • Cross-source verification: Do independent records broadly agree?

Large social or mobility datasets are not automatically representative. Data volume cannot correct selection, validity, processing, or interpretation problems, as a peer-reviewed review of bias and methodological pitfalls in social data explains. Treat these signals as supplements whose coverage and relevance must be assessed, not as a census of offline behavior.

Maintain a contradiction log containing evidence against the preferred location. To reduce survivorship bias, include closed, failed, or underperforming sites in benchmarks where reliable records are available—not only successful locations. Then vary weights, catchment boundaries, vintages, and uncertain estimates to determine whether the ranking holds.

Inspect finalists with the same field-review checklist:

  • visibility and signage constraints;
  • ingress and egress;
  • parking capacity and layout;
  • transit and pedestrian access;
  • neighboring and complementary uses;
  • utilities and infrastructure;
  • loading, storage, and physical configuration;
  • operational restrictions or site-specific constraints.

Run a separate due-diligence review for rent, buildout, labor, utilities, taxes, zoning, incentives, and break-even requirements. ACS and CBP do not answer those property-level questions. If field observations or stakeholder judgment override the quantitative ranking, require a written reason and identify the supporting evidence.

Publish an auditable site-comparison dataset

Prepare a clean public or synthetic workbook with separate tables for:

  1. candidate scores;
  2. metric definitions;
  3. source provenance;
  4. assumptions and transformations;
  5. unresolved data gaps.

Include fields for candidate, trade-area method, metric, raw measure, normalized measure, weight, score, source URL, data vintage, margin of error, status, and analyst notes. Add a data dictionary defining every variable, unit, transformation, preferred direction, and missing-value convention. Distinguish zero, unavailable, suppressed, not applicable, and not yet collected; they do not mean the same thing.

Publishing the finished spreadsheet as a public interactive data page gives stakeholders a browser-based version of the dataset rather than only a static presentation. TablePage publishes spreadsheets as public interactive data pages. Avoid documenting exact product-interface steps unless the current workflow has been verified from first-party documentation.

Use only aggregated public or synthetic records in a public example. Do not publish customer-level coordinates, confidential sales, personal information, lease terms, or other sensitive records. Keep restricted evidence in a controlled internal environment and publish only safe summaries.

The final action sequence is straightforward: define the decision, choose defensible trade areas, assemble aligned data, publish the scorecard and assumptions, stress-test the ranking, and inspect finalists before committing. The strongest result is not the site with the highest unexplained score, but the candidate that remains viable across data-quality checks, alternative assumptions, financial review, and field evidence.