AI-Assisted Transcription of Indigenous Affairs Historical Records
Accessibility · Research & Development
What it collects that can identify you
- The RG10 historical records contain personal information about individuals — including Indigenous people — documented in administrative records of the Department of Indian Affairs and Northern Development. The register confirms the system involves personal information.
Also collects operational data, which is anonymized data.
- Run by
- Library and Archives Canada (LAC)
- Where
- No fixed location
- Kept
- Not stated by the Helpful Places.
- Shared with
- Accountable organization, Me
What it is for
This system uses AI-powered handwritten text recognition to transcribe approximately six million pages of historical records from the Department of Indian Affairs and Northern Development (the RG10 collection), making them searchable and accessible to the public. It was developed by Library and Archives Canada in partnership with the Canadian Research Knowledge Network and the European company READ-COOP (Transkribus). Because the records document Indigenous history and contain personal information, access and accuracy have significant implications for affected communities. The system has since been retired.
What it collects and what happens to it
Data taken in
- The RG10 historical records contain personal information about individuals — including Indigenous people — documented in administrative records of the Department of Indian Affairs and Northern Development. The register confirms the system involves personal information.
- The primary input is approximately six million scanned page images from the RG10 Collection of Department of Indian Affairs and Northern Development historical records — administrative documents spanning colonial-era Indian Affairs governance.
Processing
- The system employs Intelligent Character Recognition (ICR) and Handwritten Text Recognition (HTR) — machine learning techniques trained to read and transcribe handwritten and printed text from scanned historical documents, with large language model assistance for contextual disambiguation.
What it does
- The system perceives and interprets scanned document images using Intelligent Character Recognition (ICR) and Handwritten Text Recognition (HTR), converting visual representations of historical handwritten and printed text into machine-readable transcriptions. Human archivists and researchers review and validate the outputs.
- Large language model capabilities are applied to improve transcription accuracy and contextual understanding of historical text, enabling better interpretation of ambiguous handwriting, archaic language, and domain-specific terminology in the Indigenous affairs records.
Outputs
- The system produces machine-readable text transcriptions of handwritten and printed historical document pages. These transcriptions make the RG10 collection searchable and accessible to members of the public, enabling research into Indigenous history.
Run by
- Library and Archives Canada (LAC) is the federal institution that deployed this transcription system in collaboration with the Canadian Research Knowledge Network (CRKN) to improve access to the RG10 historical records collection.
Government of Canada Algorithmic Impact Register — RG10 Transcription
Built by
- READ-COOP, based in Innsbruck, Austria, is the European cooperative that builds and supplies the Transkribus platform — an AI-powered handwritten text recognition tool originating from the EU-funded READ project at the University of Innsbruck. It is the technology vendor for this system.
Government of Canada Algorithmic Impact Register — RG10 Transcription
Kept for
Not stated by the Helpful Places.
Shared with
- Library and Archives Canada retains access to the transcription outputs as part of the RG10 collection management and archival mandate.
- Members of the public, listed as the primary users of the system, can access the transcribed records from the RG10 collection to support research — particularly research relevant to Indigenous history.
Stored
Not stated by the Helpful Places.
How to read the colours
Can it identify you?
- Anonymized data
- Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
- Pseudonymous data
- Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
- Identifiable data
- The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.
Who completes the loop?
- Human decides
- This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
- Human executes
- This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
- Autonomous
- This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.
Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.
- AI registerGovernment of Canada Algorithmic Impact Register — RG10 Transcription (ICR/LLM with Transkribus)Library and Archives Canada, AI Register ID: 2526-LAC-BAC-006
- AI registerGovernment of Canada Algorithmic Impact Register — RG10 Transcription
- AI registerGovernment of Canada Algorithmic Impact Register — RG10 Transcription
- Register entryPublished by the Helpful Places. Reference 68ae7918. This disclosure was drafted with AI assistance.Schema: ai@2026-05-06-beta
What you can do
Ask about this system
Questions go to the Helpful Places, not the vendor.
Your rights
- Right to Be Informed of AI UseMembers of the public accessing the RG10 collection through Library and Archives Canada are informed via the register that AI transcription technology (Transkribus/ICR/LLM) was used to produce the searchable text. The system has been retired, so this right applies to understanding historical AI use. Contact Library and Archives Canada for further information.
- Right to Correct Your DataIf you believe a transcription of a record containing your personal information or information about your family or community contains errors, you may contact Library and Archives Canada to request a correction or to flag the inaccuracy. The register confirms that personal information is involved in this system.
Risks and safeguards
- Reputational harmOCR and HTR transcription errors in historical documents involving personal information about Indigenous individuals could misrepresent names, events, or facts, causing harm to individuals or communities in the historical record. The involvement of the Canadian Research Knowledge Network and the collaborative development process with archival experts provides some mitigation through domain-specific model training, but no specific error-correction or review process is described in the register entry.
- Societal & cultural harmThe RG10 collection documents colonial administration of Indigenous peoples in Canada; inaccurate transcriptions could distort the historical record of Indigenous history, communities, and individuals. Collaboration with CRKN and the development of domain-specific recognition models is intended to improve accuracy. However, the register does not describe Indigenous community consultation in the transcription process, which is a significant gap given the cultural sensitivity of the records.