Skip to content
This is NOT an official site of the Government of Canada. Click here for the official AI registry.

AI-Assisted Technical Documentation Exploration for Statistical Analysis

Inform · Research & Development

What it collects that can identify you

About behaviour
Pseudonymous data
  • Queries and prompts entered by GC employees as they search for information within technical documentation. These interactions record what employees asked and which documents they explored.

Also collects operational data, which is anonymized data.

Run by
Statistics Canada (StatCan)
Where
No fixed location
Kept
Not stated by the Helpful Places.
Shared with
Accountable organization, Vendor

What it is for

This system uses a large language model (LLM) developed by OpenAI to help Statistics Canada employees search and navigate technical documentation more efficiently. It is designed to speed up the process of finding targeted information to support statistical analysis and processing work. The system is currently under development and is used exclusively by Government of Canada employees, not by the general public.

What it collects and what happens to it

Data taken in

Operational data
Anonymized data
  • Technical documentation — manuals, methodological guides, processing specifications, and other reference materials — used as the knowledge base for the LLM-based exploration system.
About behaviour
Pseudonymous data
  • Queries and prompts entered by GC employees as they search for information within technical documentation. These interactions record what employees asked and which documents they explored.

Processing

Language Models
  • An OpenAI large language model (LLM) is used to interpret employee queries against technical documentation and generate or retrieve targeted responses to support statistical analysis and processing work.
Search & Retrieval
  • The system retrieves targeted passages and information from technical documentation corpora in response to employee queries, using semantic search or retrieval-augmented generation (RAG) patterns.

What it does

Understanding (Semantic AI)
Human decides
  • The system uses semantic understanding to retrieve and surface relevant passages from technical documentation in response to employee queries. Human employees review and decide which information to act on.
Creating (Generative AI)
Human decides
  • The LLM component may generate natural-language responses, summaries, or synthesized answers from technical documentation. Employees review all generated content before any downstream use.

Outputs

Generated content
Anonymized data
  • Natural-language summaries, synthesized answers, and targeted excerpts generated by the LLM from technical documentation, returned to GC employees to support their analysis and processing tasks.

Run by

Statistics Canada (StatCan)
  • Statistics Canada is the federal department deploying this LLM-based tool to support internal technical documentation exploration by its employees.

Government of Canada AI Register — 2526-StatCan-001

Built by

OpenAI
  • OpenAI is the vendor that developed and supplies the large language model technology underlying this system.

Government of Canada AI Register — 2526-StatCan-001

Kept for

Not stated by the Helpful Places.

Shared with

Available to the accountable organization
  • Outputs and interaction logs are available to Statistics Canada as the deploying organization for the purposes of system monitoring, improvement, and governance oversight.
Available to vendor
  • As an OpenAI-powered system, query data and prompts may be processed by OpenAI infrastructure. The extent of vendor data access and retention under any API agreement is not specified in the register entry.

Stored

Not stated by the Helpful Places.

How to read the colours

Can it identify you?

Anonymized data
Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
Pseudonymous data
Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
Identifiable data
The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.

Who completes the loop?

Human decides
This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
Human executes
This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
Autonomous
This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.

Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.

What you can do

Ask about this system

Questions go to the Helpful Places, not the vendor.

Your rights

  • Right to Be Informed of AI UseGC employees using this system are informed they are interacting with an AI system powered by OpenAI's large language model technology. The system is identified as an AI tool in its deployment context.
  • Right to Algorithmic TransparencyAs an internal government tool, employees may request information about how the LLM-based system processes queries and generates responses through their department's AI governance and access-to-information channels.

Risks and safeguards

  • Societal & cultural harmLLM-generated responses may contain hallucinations or inaccuracies that could mislead statistical analysis and processing decisions.Safeguard: The system is currently in development and limited to GC employees who are subject-matter experts and are expected to critically review all LLM outputs before acting on them. Human oversight is built into the workflow.
  • Loss of autonomyOver-reliance on LLM-generated responses could erode employees' critical assessment of technical documentation.Safeguard: The system is positioned as an exploration and expediting tool, not a decision-making authority. Employees retain full discretion over how they use the information surfaced by the system.