AI-Assisted Technical Documentation Exploration for Statistical Analysis
Inform · Research & Development
What it collects that can identify you
- Queries and prompts entered by GC employees as they search for information within technical documentation. These interactions record what employees asked and which documents they explored.
Also collects operational data, which is anonymized data.
- Run by
- Statistics Canada (StatCan)
- Where
- No fixed location
- Kept
- Not stated by the Helpful Places.
- Shared with
- Accountable organization, Vendor
What it is for
This system uses a large language model (LLM) developed by OpenAI to help Statistics Canada employees search and navigate technical documentation more efficiently. It is designed to speed up the process of finding targeted information to support statistical analysis and processing work. The system is currently under development and is used exclusively by Government of Canada employees, not by the general public.
What it collects and what happens to it
Data taken in
- Technical documentation — manuals, methodological guides, processing specifications, and other reference materials — used as the knowledge base for the LLM-based exploration system.
- Queries and prompts entered by GC employees as they search for information within technical documentation. These interactions record what employees asked and which documents they explored.
Processing
- An OpenAI large language model (LLM) is used to interpret employee queries against technical documentation and generate or retrieve targeted responses to support statistical analysis and processing work.
- The system retrieves targeted passages and information from technical documentation corpora in response to employee queries, using semantic search or retrieval-augmented generation (RAG) patterns.
What it does
- The system uses semantic understanding to retrieve and surface relevant passages from technical documentation in response to employee queries. Human employees review and decide which information to act on.
- The LLM component may generate natural-language responses, summaries, or synthesized answers from technical documentation. Employees review all generated content before any downstream use.
Outputs
- Natural-language summaries, synthesized answers, and targeted excerpts generated by the LLM from technical documentation, returned to GC employees to support their analysis and processing tasks.
Run by
- Statistics Canada is the federal department deploying this LLM-based tool to support internal technical documentation exploration by its employees.
Built by
- OpenAI is the vendor that developed and supplies the large language model technology underlying this system.
Kept for
Not stated by the Helpful Places.
Shared with
- Outputs and interaction logs are available to Statistics Canada as the deploying organization for the purposes of system monitoring, improvement, and governance oversight.
- As an OpenAI-powered system, query data and prompts may be processed by OpenAI infrastructure. The extent of vendor data access and retention under any API agreement is not specified in the register entry.
Stored
Not stated by the Helpful Places.
How to read the colours
Can it identify you?
- Anonymized data
- Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
- Pseudonymous data
- Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
- Identifiable data
- The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.
Who completes the loop?
- Human decides
- This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
- Human executes
- This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
- Autonomous
- This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.
Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.
- AI registerGovernment of Canada Algorithmic Impact Assessment Register — OpenAI (2526-StatCan-001)Statistics Canada AI Register Entry 2526-StatCan-001, accessed 2026-05-08.
- AI registerGovernment of Canada AI Register — 2526-StatCan-001
- AI registerGovernment of Canada AI Register — 2526-StatCan-001
- Register entryPublished by the Helpful Places. Reference 277bd330. This disclosure was drafted with AI assistance.Schema: ai@2026-05-06-beta
What you can do
Ask about this system
Questions go to the Helpful Places, not the vendor.
Your rights
- Right to Be Informed of AI UseGC employees using this system are informed they are interacting with an AI system powered by OpenAI's large language model technology. The system is identified as an AI tool in its deployment context.
- Right to Algorithmic TransparencyAs an internal government tool, employees may request information about how the LLM-based system processes queries and generates responses through their department's AI governance and access-to-information channels.
Risks and safeguards
- Societal & cultural harmLLM-generated responses may contain hallucinations or inaccuracies that could mislead statistical analysis and processing decisions.Safeguard: The system is currently in development and limited to GC employees who are subject-matter experts and are expected to critically review all LLM outputs before acting on them. Human oversight is built into the workflow.
- Loss of autonomyOver-reliance on LLM-generated responses could erode employees' critical assessment of technical documentation.Safeguard: The system is positioned as an exploration and expediting tool, not a decision-making authority. Employees retain full discretion over how they use the information surfaced by the system.