Ijam
Arabic Document Intelligence. Fully offline, on your own servers.
In Arabic, iʿjām is the dotting of similar letters to remove ambiguity; 1,300 years ago it let readers read without error, and today it lets machines read Arabic documents without error. Ijam turns printed, scanned and photographed Arabic documents into searchable, archivable text on the organisation’s own machines, and no document ever leaves its network. Every word the system is unsure of goes to a human reviewer before the document is approved.
Who it’s for
Education
Student records, certificates and decisions, and search across administrative archives.
Courts, prosecution & law firms
Search across rulings and minutes, version comparison and certificates of authenticity.
Archives & records
Classification schemes, retention periods, lawful destruction and migration packages.
Manuscripts & history
Printed Ottoman script, and a names index linking Latin spellings to Arabic names.
Government
Correspondence, decisions and statutory deadlines, aligned with cybersecurity controls.
Large companies
Contracts and invoices, amount-in-words checks and tampering indicators.
How it works
Intake & quality check
A quality score per page, automatic deskew and a clear reason to rescan.
Clean-up & routing
Coloured stamps removed and lighting evened, then each page routed to the right engine for its type, with the reason shown.
Reading & Arabic correction
Text is read, then systematic Arabic errors and dot errors are corrected with a 112,000-word dictionary.
Human review
Low-confidence words go to a reviewer, and every word records which engine produced it.
Approval & archiving
Arabic search, data extraction and export to searchable PDF, Excel and ALTO.
Key features
Offline by design
Runs in a container on the organisation’s servers and installs from removable media. No data leaves the network.
Arabic & Saudi understanding
Amounts in words, Hijri dates and the Umm al-Qura calendar, personal names and dot-error correction.
PDFs & tables
Full-accuracy reading of PDF text layers, hidden-text detection and tables exported to right-to-left Excel.
Ask the archive in Arabic
Answers from a local language model with a mandatory source for every answer and prompt-injection protection.
Archive intelligence
Amount-in-words checks, conflicting-number detection, correspondence chains, a people index and deadline extraction.
Government records management
Four classification levels, Hijri retention periods, two-person destruction and PDF/A-2b export.
Enterprise security
AES-256-GCM encryption, two-factor sign-in, Active Directory and single sign-on, and a tamper-evident audit log.
Real redaction & authenticity
Redaction burned into the image and removed from the text, and a certificate that exposes any altered copy.
Learns from your corrections
A learned dictionary, training-data export and fine-tuning on the organisation’s own fonts.
We are looking for one organisation for a first 6–8 week pilot on its real documents: baseline measurement, tuning, live use with its staff, then a report with the numbers. The data always belongs to the organisation.
Measured results
Figures are measured on realistic synthetic documents and small held-out sets, not yet on a real organisation’s documents. The controls mapping is self-assessed and not yet externally audited.
Technical requirements
- Linux server (Ubuntu 22.04 or later) with Docker and Docker Compose.
- At least 16 GB RAM, 32 GB when running the local language model.
- Storage of roughly twice the archive size, with full-disk encryption.
- No GPU needed for core reading.
Request a demo of Ijam
Tell us about your organisation and needs, and we will arrange a demo with you.