Skip to content
This is NOT an official site of the Government of Canada. Click here for the official AI registry.

AI-Assisted Big Data Processing for Statistical Surveys

Planning & Decision-making · Research & Development

What it collects that can identify you

Sensitive personal information
Pseudonymous data
  • Financial institution data and administrative data from other government institutions may contain records linked to individuals or businesses. Statistics Canada applies statistical confidentiality obligations; data is processed under strict access controls and linkage is typically performed with tokenized identifiers.

Also collects operational data and about a measurement, which is anonymized data.

Run by
Statistics Canada (StatCan)
Where
No fixed location
Kept
Not stated by the Helpful Places.
Shared with
Accountable organization
Your copy
You cannot see the data it holds about you. What you can do

What it is for

Statistics Canada uses machine learning models to process large volumes of administrative and commercial data — such as retail scanner records, satellite images, and financial transactions — in order to produce national statistics. This system reduces the burden on businesses and individuals who would otherwise need to respond to surveys. The outputs are statistical aggregates used internally by government employees, not decisions made about specific individuals.

What it collects and what happens to it

Data taken in

Operational data
Anonymized data
  • Retail scanner data, large transport transactional data files, housing data, international trade transactions, and administrative data from other government institutions. These are large-scale operational records not typically linked to named individuals.
About a measurement
Anonymized data
  • Satellite images processed as pixel-level measurements of land use, agricultural conditions, or infrastructure. These inputs do not identify individuals.
Sensitive personal information
Pseudonymous data
  • Financial institution data and administrative data from other government institutions may contain records linked to individuals or businesses. Statistics Canada applies statistical confidentiality obligations; data is processed under strict access controls and linkage is typically performed with tokenized identifiers.

Processing

Classification & Prediction
  • Machine learning classifiers and predictive models are used to assign statistical codes, estimate economic indicators, and derive statistical variables from administrative and commercial data sources, replacing traditional survey collection.
Computer Vision
  • Computer vision techniques are applied to satellite images to extract statistical signals such as land cover classification, agricultural yield estimates, or construction activity indicators.

What it does

Deciding (Analytical AI)
Human decides
  • Models classify, predict, and score inputs derived from big data sources (retail scanner data, financial data, administrative data) to generate statistical estimates. Government employees (primary users) review and act on these analytical outputs.
Sensing (Perceptive AI)
Human decides
  • The system processes satellite images and other unstructured data sources, converting raw signals into structured features that feed downstream statistical models.

Outputs

Operational data
Anonymized data
  • Statistical estimates, indicators, and derived variables produced for use by Statistics Canada staff in compiling national accounts, price indices, and other official statistics. Outputs are aggregated and not linked to identifiable individuals.
A recommendation or prediction
Anonymized data
  • Model outputs serve as recommendations or predictions that GC employees use to make final statistical determinations. Human review is a standard part of the statistical production workflow.

Run by

Statistics Canada (StatCan)
  • Statistics Canada is the federal department responsible for producing national statistics. It deploys and operates these machine learning models to process big data as part of its statistical programs.

GC AI Register — 2526-StatCan-008

Built by

Government of Canada
  • The system was developed internally by the Government of Canada, with Statistics Canada as the developing department.

GC AI Register — 2526-StatCan-008

Kept for

Not stated by the Helpful Places.

Shared with

Available to the accountable organization
  • Outputs are available to Statistics Canada employees (GC employees) who are the primary users of the system for statistical production purposes.
Not available to me
  • Individual members of the public do not have access to the model outputs or the underlying data processed by this system. Published statistics derived from these models are made publicly available through Statistics Canada's standard dissemination channels, but individual-level model outputs are not accessible.

Stored

Not stated by the Helpful Places.

How to read the colours

Can it identify you?

Anonymized data
Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
Pseudonymous data
Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
Identifiable data
The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.

Who completes the loop?

Human decides
This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
Human executes
This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
Autonomous
This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.

Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.

What you can do

Ask about this system

Questions go to the Helpful Places, not the vendor.

Your rights

  • Right to Algorithmic TransparencyThis system is listed on the Government of Canada's Algorithmic Impact Assessment register, which provides public disclosure of its purpose, data sources, and deployment status. Further information about Statistics Canada's methods is published in methodological guides available at statcan.gc.ca.

Risks and safeguards

  • Societal & cultural harmMachine learning models trained on administrative and commercial data may introduce systematic biases into national statistics, potentially misrepresenting certain populations or economic sectors.Safeguard: Statistics Canada applies rigorous quality assurance and validation processes before integrating model outputs into official statistics. The use of ML is intended to supplement, not replace, methodological expertise, and outputs undergo review by GC employees.