AI-Assisted Extraction of Immigration Records from Historical Documents
Research & Development
What it collects that can identify you
- Scanned pages from 11,231 pages of the Canada Gazette containing historical immigration records, which include personal information about named individuals (immigrants). These are historical records with identifiable personal data such as names and immigration details.
- Run by
- Library and Archives Canada (LAC)
- Where
- No fixed location
- Kept
- Not stated by the Helpful Places.
- Shared with
- Accountable organization, Me
What it is for
This system used Amazon Web Services Textract to automatically read and index 11,231 pages of historical Canada Gazette immigration records, making them searchable for genealogy research. It extracted structured text and table data from scanned documents and indexed the results using ElasticSearch. The system involved personal information drawn from historical immigration records and was developed by Library and Archives Canada using a vendor-supplied AI service. This system has been retired and is no longer in active operation.
What it collects and what happens to it
Data taken in
- Scanned pages from 11,231 pages of the Canada Gazette containing historical immigration records, which include personal information about named individuals (immigrants). These are historical records with identifiable personal data such as names and immigration details.
Processing
- AWS Textract applies computer vision to interpret scanned document images — reading printed text, detecting table structures, and extracting fields from historical immigration record pages.
- ElasticSearch indexing of the extracted text enables users to search across the 11,231-page corpus of Canada Gazette immigration records to find relevant entries for genealogy research.
What it does
- AWS Textract performs OCR and structured text extraction from scanned document pages, turning raw page images into machine-readable text and table data. Humans then use the indexed output for genealogy research queries.
- ElasticSearch indexing classifies and ranks extracted text records to support full-text search queries, enabling researchers to find relevant immigration entries across the document corpus.
Outputs
- Indexed and searchable text records extracted from historical immigration entries in the Canada Gazette — containing names and other personal details of historical immigrants, surfaced in response to genealogy research queries.
Run by
- Library and Archives Canada is the federal government institution that developed and deployed this system to extract and index immigration records from the Canada Gazette for genealogy research purposes.
Built by
- Amazon Web Services supplied the AWS Textract AI service used to perform optical character recognition and text extraction. The system was developed by Library and Archives Canada using this vendor-provided technology.
Kept for
Not stated by the Helpful Places.
Shared with
- The indexed immigration records were accessible to Library and Archives Canada employees as primary users of the system.
- The indexed records were also made available to the general public for genealogy research purposes, allowing individuals to search for their own or ancestors' immigration records.
Stored
Not stated by the Helpful Places.
How to read the colours
Can it identify you?
- Anonymized data
- Data about people with the link to who is broken. Stripped of identifiers, blurred, aggregated, or noised so this system can’t reasonably tie a record back to an individual.
- Pseudonymous data
- Each person’s data is tied to a token (hash, ID, template) that lets this system recognise the same person across events, but the token itself doesn’t reveal a name. Reidentification is possible with extra information.
- Identifiable data
- The data either contains a direct identifier (name, address, account name, recognisable face or voice, plate number) or carries a token this system uses to look up legal identity during processing.
Who completes the loop?
- Human decides
- This mode suggests; a person decides what to do next. The AI is always advisory — a human is in the loop on every decision. Example: a triage tool ranks cases for a clinician who chooses which to see first.
- Human executes
- This mode decides; a person carries out the result. Example: an optimizer plans the day’s trash-collection routes, and drivers run them.
- Autonomous
- This mode decides and acts on its own. No person reviews each decision or carries out the resulting action.
Definitions from the DTPR standard. Amber is about your data, violet about who decides. The fuller the shape and the deeper the colour, the more identifying the data or the less a person is involved.
- AI registerCanada AI and Data Use Register — Canada Gazette Immigration Records AWS Textract Project (2526-LAC-BAC-004)Library and Archives Canada, Government of Canada Algorithmic Impact Assessment Register, entry 2526-LAC-BAC-004.
- AI registerCanada AI Register entry 2526-LAC-BAC-004
- AI registerCanada AI Register entry 2526-LAC-BAC-004
- Register entryPublished by the Helpful Places. Reference eafd6c56. This disclosure was drafted with AI assistance.Schema: ai@2026-05-06-beta
What you can do
Ask about this system
Questions go to the Helpful Places, not the vendor.
Your rights
- Right to Algorithmic TransparencyMembers of the public may request information about how this AI system processes historical immigration records by contacting Library and Archives Canada. Details about the system have been disclosed in the Government of Canada AI and Data Use Register.
- Right to AccessIndividuals whose personal information appears in the Canada Gazette immigration records may contact Library and Archives Canada to request access to their data, consistent with the federal Privacy Act and Access to Information Act.
Risks and safeguards
No risks or safeguards have been published for this system yet.