Skip to content
This is NOT an official site of the Government of Canada. Click here for the official AI registry.

AI-Assisted Survey Response Classification for Statistics Canada

Employment & Work · Research & Development

What it collects

About behaviour
Anonymized data
  • Past survey responses from Statistics Canada statistical programs. These are text responses collected in prior survey cycles and used to train or calibrate the classification model. The register confirms no personal information is involved.
Operational data
Anonymized data
  • Classification standards used by Statistics Canada statistical programs. These are reference documents that define the official categories and codes the system must assign to incoming text responses.
Run by
Statistics Canada (StatCan)
Where
No fixed location
Kept
Not stated by the Helpful Places.
Shared with
Accountable organization
Your copy
You cannot see the data it holds about you. What you can do

What it is for

This system uses a large language model (LLM) to automatically code and classify text responses from Statistics Canada statistical programs. It is designed to assist government employees who process survey data, not members of the public. The system does not handle personal information, and users are informed when AI is involved in the process.

What it collects and what happens to it

Data taken in

About behaviour
Anonymized data
  • Past survey responses from Statistics Canada statistical programs. These are text responses collected in prior survey cycles and used to train or calibrate the classification model. The register confirms no personal information is involved.
Operational data
Anonymized data
  • Classification standards used by Statistics Canada statistical programs. These are reference documents that define the official categories and codes the system must assign to incoming text responses.

Processing

Language Models
  • A large language model (LLM) is used to read and interpret text responses and assign classification codes. The system explores horizontal scaling approaches for this LLM-based text coding task. The specific model architecture or provider is not disclosed in the register entry.
Classification & Prediction
  • The system assigns classification labels to text inputs. Classification is based on patterns learned from past survey responses mapped against official classification standards.

What it does

Understanding (Semantic AI)
Human decides
  • The LLM reads text responses from statistical surveys and maps them to classification codes based on classification standards. It understands the meaning of the text to determine the appropriate category. Government employees review and use the results.
Deciding (Analytical AI)
Human decides
  • The system classifies and assigns category labels to text inputs based on learned patterns from past survey responses and classification standards. Human employees review the assigned codes before they are used.

Outputs

A recommendation or prediction
Anonymized data
  • The system outputs classification codes or labels for each text response. These are recommendations that government employees use in their statistical coding work. The register indicates AI use is disclosed to users, suggesting human review remains in the loop.

Run by

Statistics Canada (StatCan)
  • Statistics Canada is the federal department responsible for producing statistics on Canada's economy, society, and population. It is developing and deploying this LLM-based coding tool for use by its employees in statistical classification work.

Government of Canada AI Register — 2526-StatCan-021

Built by

Government of Canada
  • The system was developed internally by the Government of Canada. No external vendor is identified in the register entry.

Government of Canada AI Register — 2526-StatCan-021

Kept for

Not stated by the Helpful Places.

Shared with

Available to the accountable organization
  • Outputs (classification codes) are available to Statistics Canada employees who use the system in their statistical coding workflows. The register identifies GC employees as primary users.
Not available to me
  • Members of the public (including survey respondents) do not have access to the system's outputs or internal coding decisions. The system is for internal government use only.

Stored

Not stated by the Helpful Places.

How to read the colours

Can it identify you?

Anonymized data
Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
Pseudonymous data
Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
Identifiable data
The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.

Who completes the loop?

Human decides
This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
Human executes
This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
Autonomous
This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.

Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.

What you can do

Ask about this system

Questions go to the Helpful Places, not the vendor.

Your rights

  • Right to Be Informed of AI UseStatistics Canada discloses to users (Government of Canada employees) when AI is being used in the classification process. This right applies to the internal users of the system, not to survey respondents.
  • Right to Algorithmic TransparencyAs a system listed on the Government of Canada's public AI register, information about how the LLM classification tool operates is made available. Users and the public can consult the register entry for details about the system's purpose, data sources, and development status.

Risks and safeguards

  • Societal & cultural harmSystematic misclassification of survey responses could introduce bias into official Statistics Canada data products, affecting data quality and public trust in national statistics.Safeguard: The system is noted as in development, suggesting a staged rollout with testing. AI use is disclosed to the government employees who use the outputs, enabling human oversight and correction of errors before results are finalized.