Skip to content
This is NOT an official site of the Government of Canada. Click here for the official AI registry.

AI-Assisted Text Analysis for Employee Surveys

Employment & Work · Research & Development

What it collects

Operational data
Anonymized data
  • Open-ended survey responses from four CRA employee surveys, collected as internal survey data. The register states no personal information is involved, indicating responses are treated as de-identified or aggregated administrative data.
Run by
Canada Revenue Agency (CRA)
Where
No fixed location
Kept
Not stated by the Helpful Places.
Shared with
Accountable organization
Your copy
You cannot see the data it holds about you. What you can do

What it is for

This system uses natural language processing and machine learning to analyze open-ended responses from employee surveys at the Canada Revenue Agency. It identifies patterns, clusters themes, and generates predictive and prescriptive insights to support evidence-based decisions about human resources programs. The system processes internal survey data and its outputs are used by Government of Canada employees in the Data Division. No personal information is involved in the analysis.

What it collects and what happens to it

Data taken in

Operational data
Anonymized data
  • Open-ended survey responses from four CRA employee surveys, collected as internal survey data. The register states no personal information is involved, indicating responses are treated as de-identified or aggregated administrative data.

Processing

Language Models
  • Python-based natural language processing (NLP) is used to analyze free-text survey responses, enabling theme extraction, pattern detection, and sentiment-related analysis.
Clustering & Segmentation
  • Cluster analysis groups similar survey responses to reveal underlying themes and patterns within the employee survey data without pre-defined categories.
Classification & Prediction
  • Predictive and prescriptive models are applied to survey response data to forecast trends and support program planning decisions within the CRA's human resources branch.

What it does

Deciding (Analytical AI)
Human decides
  • The system scores, classifies, clusters, and ranks patterns in employee survey text. Outputs inform human decision-makers who determine follow-up actions for HR programs; the system does not make binding decisions autonomously.
Understanding (Semantic AI)
Human decides
  • Natural language processing is used to extract meaning from open-ended survey responses, identify themes, and link related ideas across the four employee surveys analyzed.

Outputs

A recommendation or prediction
Anonymized data
  • The system produces pattern summaries, cluster groupings, predictions, and prescriptive recommendations that inform evidence-based decisions for human resources programs. These are advisory outputs reviewed by GC employees, not automated determinations.

Run by

Canada Revenue Agency (CRA)
  • The Canada Revenue Agency (CRA) is the federal department that administers tax laws and benefit programs for the Government of Canada. CRA's Data Division deploys and operates this text analysis system for internal human resources research.

Text Analysis for HRB - Data Division Supported Surveys — Canada AI Register

Built by

Not stated by the Helpful Places.

Kept for

Not stated by the Helpful Places.

Shared with

Not available to me
  • Survey respondents cannot access the outputs or analytical results produced by this system. Access is restricted to GC employees working within the CRA Data Division for internal HR research purposes.
Available to the accountable organization
  • Analytical outputs are available to CRA's Data Division and the Human Resources Branch (HRB) for internal program planning and evidence-based decision-making.

Stored

Not stated by the Helpful Places.

How to read the colours

Can it identify you?

Anonymized data
Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
Pseudonymous data
Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
Identifiable data
The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.

Who completes the loop?

Human decides
This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
Human executes
This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
Autonomous
This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.

Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.

What you can do

Ask about this system

Questions go to the Helpful Places, not the vendor.

Your rights

  • Right to Be Informed of AI UseGC employees who participated in the four surveys analyzed by this system may not have been specifically notified that their responses would be subject to automated NLP analysis. The Canada AI Register entry for this system is publicly accessible and discloses its use.
  • Right to Algorithmic TransparencyThe Canada AI Register publicly discloses that CRA uses Python-based NLP, cluster analysis, and predictive modeling on employee survey data. Further details about the specific methods and models used are not publicly documented in the register entry.

Risks and safeguards

  • Societal & cultural harmAutomated pattern analysis of employee survey responses could surface inferred groupings or themes that inadvertently reflect or reinforce workplace biases, even without processing personal information.Safeguard: The register states no personal information is involved; outputs are reviewed by GC employees rather than used for automated decisions, maintaining human oversight of all interpretations and program actions derived from the analysis.