How we evaluate a fabrication website
The purpose of this work is not to produce another subjective website score. It is to determine, as rigorously as is reasonable, which characteristics of these websites can be objectively observed, which have credible theoretical or empirical reason to affect customer acquisition, which can be shown to affect prospective-customer behaviour, and which ultimately predict qualified inquiries, quotes, customers and revenue.
We treat it as applied human-factors, HCI, marketing and experimental research, and we publish the method so anyone can check what a given claim actually rests on.
The fundamental rule: never infer a causal customer-acquisition outcome directly from a scraped website feature.
The evidence hierarchy
Every finding sits at a level, and the level governs the language we are allowed to use about it. A claim is never promoted to a stronger level than the method that produced it.
| Level | Method | What may be claimed |
|---|---|---|
| 1 | Observed condition Something directly measurable from the website — a file upload exists or does not, a phone number is tap-to-call or is not, an LCP measured at X seconds, analytics code detectable or not. |
May be reported as fact, scoped to what was fetched. |
| 2 | Evidence-supported risk A Level 1 observation that corresponds to an established mechanism in published research or standards: cognitive burden, weak information scent, interaction cost, accessibility barriers, form friction. |
Stated as a risk, hypothesis or likely mechanism. Never as a demonstrated customer loss. |
| 3 | User-behaviour evidence Representative prospective customers attempt realistic tasks. Task completion, first-click success, errors, abandonment, ability to judge capability, RFQ completeness, perceived difficulty and trust. |
A feature may be said to affect or be associated with measured behaviour within the study. |
| 4 | Field behaviour Real analytics and inquiry data through the funnel: qualified traffic → service page → inquiry initiation → completion → quote-ready inquiry → quote issued → accepted → customer. |
Associations between site experience and acquisition behaviour may be evaluated. |
| 5 | Causal validation Randomised controlled experiments assigning visitors to existing versus modified experiences, with hypothesis, outcomes, power, stopping rules and duration fixed in advance. |
A causal claim may be made, reported with effect size and interval. |
We always distinguish observed, associated, predicted and causal. “64% of sampled sites did not provide file upload within the primary inquiry workflow” is a statement of fact. “64% of these businesses lose customers because of it” is not something a scrape can support, and we do not write it.
Three nested studies
Breadth, then mechanism, then causality. No single layer is the contribution; the combination is.
| Study | Scale | What it establishes |
|---|---|---|
| Market observatory | 1,008+ sites | What actually exists in the market |
| Human factors study | ~30–100 sites, plus target users | Which configurations help or hinder realistic customer tasks |
| Field experiments | Client sites | What actually changes qualified acquisition and shop workload |
What we measure
Customer acquisition is treated as a system of distinct constructs, reported separately. They are deliberately not combined into a single “website quality” number, because summing unlike things with invented weights produces a figure that looks precise and means nothing.
Discovery & findability
Can an appropriate prospective customer discover the business?
- Indexability, robots directives, sitemap
- Titles, meta descriptions, canonical configuration
- Service-area and locally relevant content
- Structured data
- Page architecture and internal links
We do not claim a ranking impact for any individual feature unless the evidence supports it.
Capability comprehension
Can a prospect quickly determine whether this shop can perform their work?
- Primary service immediately identifiable
- Materials, processes and equipment identified
- Project scale and geographic region understandable
- Custom versus standard work made clear
- Project examples connected to stated capabilities
Tested with scenarios: “You need 20 powder-coated steel brackets made from a drawing. Can this company produce it, and what would you do next?” We measure correct judgment, confidence, time to answer, and which information the person relied on — which turns “clear messaging” into task performance.
Information scent & navigation
Do labels, headings and links signal that the wanted information lies behind them?
- First-click testing
- Task-based navigation and path analysis
- Misclicks, backtracking, abandonment
Grounded in Information Foraging Theory. The existence of navigation is not evidence of effective navigation.
Credibility & trust
What observable signals of legitimacy does the site carry?
- Real company identity, location or service area
- Named people, original project photography
- Identifiable prior projects, testimonials, third-party reviews
- Licences and certifications where relevant
- Visible errors, broken elements, content recency
Observable credibility cues are coded separately from perceived credibility, because perception has to be measured with people rather than inferred from markup.
RFQ & contact effectiveness
Can an inquiry be submitted, and is it useful when it arrives?
- File and photo attachment, supported types
- Field labelling, required versus optional, error handling
- Mobile completion, confirmation, expected response time
- Save and resume where appropriate
This is where generic conversion advice fails fabrication. “Fewer fields is better” assumes every question is friction — but a two-field “name and message” form is trivial for the customer and expensive for the shop.
Usability
Effectiveness, efficiency and satisfaction in a specified context of use.
- Structured heuristic evaluation
- Realistic task-based testing with participants
- A validated post-task instrument such as SUS or UMUX-LITE
Framed by ISO 9241-11:2018. Heuristic inspection supplements participant testing; it does not replace it, and an LLM does not substitute for either.
Accessibility
Does the site meet WCAG 2.2 where technically assessable?
- Semantic headings and text alternatives
- Keyboard access and focus behaviour
- Labels, contrast, target sizing
- Form errors and instructions
- Responsive and mobile interaction
Failures are reported as accessibility findings. We do not convert an accessibility defect into an acquisition-loss percentage.
Technical experience
Standardised measurement instead of “the site feels slow”.
- Largest Contentful Paint
- Interaction to Next Paint
- Cumulative Layout Shift
- Image payload, broken resources, HTTPS
Laboratory measurement and real-user field measurement are recorded separately, because they are not interchangeable.
Measurement readiness
Does the business appear able to observe its own funnel?
- Analytics script detectable
- Tag manager detectable
- Conversion events detectable where observable
- Phone tracking, CRM integration, UTM capture
A public scrape cannot establish that a company has no analytics. It establishes only that no client-side implementation was detected by our method. Server-side analytics, hosting analytics, call tracking and CRM reporting are not externally visible.
Quote readiness versus customer effort
This is the part of the method that is specific to custom production, and the part we think matters most. Two constructs are measured against each other rather than one being assumed: customer effort, how hard the inquiry is to submit, and quote readiness, whether the inquiry contains enough for the shop to take the next meaningful action.
For each trade we build a Minimum Quote Information Set from interviews with people who actually quote work — asking about the last five inquiries that were difficult to price and what was missing, rather than what they would like on their website. Then:
RFQ Completeness = required quote information collected ÷ required information in that trade’s Minimum Quote Information Set
A machine shop, a cabinetmaker, a sign shop, a screen printer and a remodelling contractor are not judged against the same intake specification, because they cannot quote from the same information. The objective is not the shortest form. It is the minimum customer burden sufficient to create a useful inquiry.
How the large sample was collected
The market observatory covers 1,008 publicly accessible trade and fabrication websites, measured in 2026. Businesses were discovered by structured web search across metropolitan area and trade, with national chains, multi-city template networks and directory aggregators excluded. Queries, cities, trades and collection dates are recorded.
Search-discovered businesses are not a representative sample of the industry. Discoverability itself determines inclusion — and discoverability is one of the things being studied. This is a structured sample of publicly discoverable trade-business websites, not all trade businesses.
How the measurements are validated
Automated coding is never treated as ground truth. A stratified subset is coded independently by people using the same codebook, and machine output is compared against that reference using precision, recall, F1, and false-positive and false-negative rates. For subjective categories at least two independent coders are used and agreement is reported with Cohen’s κ. Disagreement is investigated rather than averaged away.
Any variable whose reliability is inadequate is redefined, recoded by hand, reported with its uncertainty, or removed. Automated extraction operates at a scale people cannot; people validate the instrument. Those are different jobs and we do not swap them.
Every variable carries an explicit “cannot determine reliably” value. A measurement that silently records absence when the honest answer is uncertainty will overstate every finding built on it.
What the numbers are for
The measures that matter are the shop’s economics, not website engagement. We do not optimise page views, time on page, or raw form volume when they conflict with better business outcomes.
Clarification burden deserves particular attention. An intake system can create real value with no increase in inquiry volume at all: the same number of inquiries, carrying more complete information, requiring fewer exchanges before a price can be given, freeing quoting capacity. That is acquisition efficiency, and it is invisible to any measure that counts submissions.
How changes are tested
Wherever traffic permits we test individual hypotheses with randomised controlled experiments, rather than replacing an entire website and attributing everything that follows to “the redesign”. Hypothesis, primary outcome, secondary outcomes, sample size, stopping rules and duration are fixed before analysis. We report effect size and interval, not only whether p < .05 — a statistically detectable improvement can still be commercially trivial.
Where randomisation is impossible we use interrupted time-series, difference-in-differences, staggered implementation or matched comparison sites, and state plainly that the causal certainty is weaker and what the assumptions are.
The method is designed so hypotheses can be rejected. Finding that a presumed best practice has no meaningful relationship with qualified customer acquisition is an equally successful result, and we would rather publish that than a dramatic one.
References
Primary literature and standards, in preference to marketing sources.
- ISO 9241-11:2018. Ergonomics of human-system interaction — Part 11: Usability: definitions and concepts.
- Nielsen, J. & Molich, R. (1990). Heuristic evaluation of user interfaces. CHI ’90. doi:10.1145/97243.97281
- Fogg, B.J. et al.. Stanford Web Credibility Project — guidelines derived from research involving over 4,500 participants.
- Pirolli, P. & Card, S.. Information Foraging Theory — information scent and navigation behaviour.
- Lindgaard, G. et al. (2006). Attention web designers: You have 50 milliseconds to make a good first impression. Behaviour & Information Technology. doi:10.1080/01449290500330448
- Lewis, J.R., Utesch, B.S. & Maher, D.E. (2013). UMUX-LITE: when there’s no time for the SUS. CHI ’13. doi:10.1145/2470654.2481287
- Sauro, J. & Lewis, J.R. (2005). Estimating completion rates from small samples using binomial confidence intervals. doi:10.1177/154193120504902407
- Cohen, J. (1960). A coefficient of agreement for nominal scales. doi:10.1177/001316446002000104
- W3C. Web Content Accessibility Guidelines (WCAG) 2.2.
- Google. Core Web Vitals — current definitions and measurement guidance (LCP, INP, CLS).
- Kohavi, R. et al.. Online controlled experiments and A/B testing.
For each criterion we record its source, the type of evidence, the proposed mechanism, the strength of that evidence, and whether the criterion is descriptive, predictive or experimentally validated. Where evidence is weak, contradictory, industry-specific, or extrapolated from e-commerce rather than custom fabrication, we say so rather than quietly borrowing its authority.