Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. How Do We Publish Sector Statistics Without Anyone Working Out an Individual Member's Figures?
Problems We Solve

How Do We Publish Sector Statistics Without Anyone Working Out an Individual Member's Figures?

Trade associations fear a published table will reveal one member's figures. We build aggregation with your disclosure rules applied before anything goes out.

Updated 3 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Disclosure risk creeps in when statistics are built by hand in spreadsheets and nobody checks every cell against your confidentiality rules. We build the aggregation step so your rules, such as minimum firms per cell and maximum share for one firm, are applied automatically, cells that break them are suppressed along with the ones that would give them away, and every published table has a record of the check.

A member rings about Table 4

The quarterly statistics went out yesterday. This morning a member's commercial director calls. Table 4 shows production in one region for one product group, and there are only two firms in that region making that product. Their competitor can subtract its own figure and see theirs. They are not happy, and they want to know how it happened.

Nobody meant it to. The analyst built the tables in Excel as always, checked the big cells, and missed that one breakdown had become sparse since a firm closed its plant. Your confidentiality undertaking to members says this should not happen. The trust that makes firms share data at all is the thing at risk.

Why sparse cells slip through

Disclosure rules are usually simple to state and tedious to apply. Applying them by eye across dozens of tables, every period, is where they fail.

  • Rules exist in a policy document, not in the spreadsheets that build the tables.
  • Membership changes (a closure, a merger, a new entrant) change which cells are sparse without anyone noticing.
  • Suppressing one cell is not enough when row and column totals let a reader work it out, so secondary suppression is needed and easy to forget.
  • Ad hoc requests from government or press get built quickly, outside the usual checks.
  • There is no record of what was checked for each table, so nobody can show the rules were applied.

Your rules, and whether they meet any legal or competition-law requirements, are for your board and advisers. What we fix is making sure the rules you have actually run on every table.

What a disclosure costs

Members stop sending data, or send less. A statistics series that took years to build loses coverage. Staff spend days on apologies, investigations and reassurance. And your team becomes nervous about publishing detail at all, which makes the statistics less useful to the members who do contribute.

Aggregation with the rules built in

We move table production into a small, repeatable pipeline that knows your confidentiality rules and applies them before any figure leaves the building.

  1. Clean member returns sit in one dataset with firm identity attached, restricted to the analysts who need it.
  2. Each published table is defined once: which measure, which breakdowns, which period.
  3. For every cell, the pipeline counts contributing firms and each firm's share, and tests them against your rules, such as a minimum number of firms or a maximum share for the largest contributor.
  4. Cells that fail are suppressed, and the pipeline then finds any other cells that would let a reader recover them through totals, and suppresses those as well, or merges categories if your rules prefer that.
  5. The analyst reviews a preview with suppressed cells marked and the reason for each, then approves.
  6. Approved tables are written out to Excel, CSV or your publication template, and a disclosure log records the rule results for every table and period.
CheckWhat it catchesResult in the published table
Minimum contributorsCells with too few firmsCell suppressed or merged
DominanceOne firm providing most of a cellCell suppressed
Secondary suppressionTotals that reveal a suppressed cellExtra cells suppressed
Period comparisonA cell that became sparse since last periodAnalyst warned before publishing
Ad hoc requestOne-off tables outside the usual setSame checks, same log

The pipeline can be written in Python or R, run on your own machine or in Azure or AWS, and your analysts can read and change it. We avoid building anything only we can maintain.

Publication day, afterwards

The analyst runs the tables, reviews the suppressions, and publishes. When a journalist or a department asks for a special breakdown, it goes through the same checks instead of being built by hand under pressure. If a member ever asks how their data is protected, you can show them the rules and the log for the table they are worried about.

Warning signs in your statistics process

  • Your confidentiality rules are written down but applied by eye.
  • Tables are built by copying formulas from last period's workbook.
  • Ad hoc requests skip the normal checks because they are urgent.
  • Nobody checks whether totals reveal a suppressed cell.
  • You could not show a member which checks were applied to a given table.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Do you set the confidentiality rules?

No. Your board and advisers set them. We encode the rules you give us and make them run on every table, every time.

Can our analysts still build custom tables?

Yes. They define a table and the pipeline applies the same checks, so custom work gets the same protection.

Does this need a big database?

No. Many sector datasets fit comfortably in a small database or even well-structured files. We size it to what you have.

What do you need from us?

Your written disclosure rules, a set of recent tables, and access to the anonymised or restricted source data under your data policy.

What drives the cost?

The number of tables and breakdowns, how complex your rules are, and whether secondary suppression is needed across many dimensions.

Keep reading

More on Problems We Solve

Start here

Tell us how your association's data and policy work gets done

Describe the survey, the benchmark report or the consultation process that eats your team's weeks, the tools involved and who signs things off. We will tell you what we would automate and what should stay with your analysts and policy staff, and if a smaller change would fix it, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →