AI-Assisted Survey Response Classification for Statistics Canada
Employment & Work · Research & Development
What it collects
- Past survey responses from Statistics Canada statistical programs. These are text responses collected in prior survey cycles and used to train or calibrate the classification model. The register confirms no personal information is involved.
- Classification standards used by Statistics Canada statistical programs. These are reference documents that define the official categories and codes the system must assign to incoming text responses.
- Run by
- Statistics Canada (StatCan)
- Where
- No fixed location
- Kept
- Not stated by the Helpful Places.
- Shared with
- Accountable organization
- Your copy
- You cannot see the data it holds about you. What you can do
What it is for
This system uses a large language model (LLM) to automatically code and classify text responses from Statistics Canada statistical programs. It is designed to assist government employees who process survey data, not members of the public. The system does not handle personal information, and users are informed when AI is involved in the process.
What it collects and what happens to it
Data taken in
- Past survey responses from Statistics Canada statistical programs. These are text responses collected in prior survey cycles and used to train or calibrate the classification model. The register confirms no personal information is involved.
- Classification standards used by Statistics Canada statistical programs. These are reference documents that define the official categories and codes the system must assign to incoming text responses.
Processing
- A large language model (LLM) is used to read and interpret text responses and assign classification codes. The system explores horizontal scaling approaches for this LLM-based text coding task. The specific model architecture or provider is not disclosed in the register entry.
- The system assigns classification labels to text inputs. Classification is based on patterns learned from past survey responses mapped against official classification standards.
What it does
- The LLM reads text responses from statistical surveys and maps them to classification codes based on classification standards. It understands the meaning of the text to determine the appropriate category. Government employees review and use the results.
- The system classifies and assigns category labels to text inputs based on learned patterns from past survey responses and classification standards. Human employees review the assigned codes before they are used.
Outputs
- The system outputs classification codes or labels for each text response. These are recommendations that government employees use in their statistical coding work. The register indicates AI use is disclosed to users, suggesting human review remains in the loop.
Run by
- Statistics Canada is the federal department responsible for producing statistics on Canada's economy, society, and population. It is developing and deploying this LLM-based coding tool for use by its employees in statistical classification work.
Built by
- The system was developed internally by the Government of Canada. No external vendor is identified in the register entry.
Kept for
Not stated by the Helpful Places.
Shared with
- Outputs (classification codes) are available to Statistics Canada employees who use the system in their statistical coding workflows. The register identifies GC employees as primary users.
- Members of the public (including survey respondents) do not have access to the system's outputs or internal coding decisions. The system is for internal government use only.
Stored
Not stated by the Helpful Places.
How to read the colours
Can it identify you?
- Anonymized data
- Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
- Pseudonymous data
- Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
- Identifiable data
- The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.
Who completes the loop?
- Human decides
- This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
- Human executes
- This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
- Autonomous
- This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.
Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.
- AI registerGovernment of Canada Algorithmic Impact Assessment Register — Using LLM for Coding/classification of StatCan statistical programsAI Register ID: 2526-StatCan-021. Statistics Canada, Government of Canada.
- AI registerGovernment of Canada AI Register — 2526-StatCan-021
- AI registerGovernment of Canada AI Register — 2526-StatCan-021
- Register entryPublished by the Helpful Places. Reference 7b7e86c4. This disclosure was drafted with AI assistance.Schema: ai@2026-05-06-beta
What you can do
Ask about this system
Questions go to the Helpful Places, not the vendor.
Your rights
- Right to Be Informed of AI UseStatistics Canada discloses to users (Government of Canada employees) when AI is being used in the classification process. This right applies to the internal users of the system, not to survey respondents.
- Right to Algorithmic TransparencyAs a system listed on the Government of Canada's public AI register, information about how the LLM classification tool operates is made available. Users and the public can consult the register entry for details about the system's purpose, data sources, and development status.
Risks and safeguards
- Societal & cultural harmSystematic misclassification of survey responses could introduce bias into official Statistics Canada data products, affecting data quality and public trust in national statistics.Safeguard: The system is noted as in development, suggesting a staged rollout with testing. AI use is disclosed to the government employees who use the outputs, enabling human oversight and correction of errors before results are finalized.