Skip to content
Onigiri Tech

Artificial Intelligence / Automation

Document Intelligence Workspace

An AI workspace that extracts structured data from incoming documents and lets staff search internal knowledge with cited answers.

  • Field extraction
  • Validation rules
  • Human review
  • Semantic search
  • Cited answers
  • Permission-aware retrieval

Context

An organization processing high volumes of forms and supporting documents, with policies spread across shared drives.

The challenge

Staff manually read each document to type the same fields into an internal system.

Finding the current version of a policy depended on knowing who to ask.

Approach

  1. 01

    Measured on real documents

    An evaluation set built from representative documents before choosing models.

  2. 02

    Extraction with review

    Fields extracted and validated; anything uncertain goes to a reviewer with the source highlighted.

  3. 03

    Grounded answers

    Search and Q&A answer only from indexed documents the user is allowed to see, with citations.

  4. 04

    Integrated, not separate

    Results flow directly into the existing case system through its API.

What we built

  • Ingestion pipeline
  • Extraction service
  • Review interface
  • Knowledge index
  • Assistant UI
  • Evaluation dashboard
  • Python
  • FastAPI
  • LLM APIs
  • PostgreSQL + pgvector
  • OCR engine
  • Next.js

Outcome

  • Routine fields extracted automatically with human review for exceptions
  • Policy questions answered with a link to the source document
  • Accuracy tracked continuously against an evaluation set

Next case study

Internal Requests & Approvals Suite

Complex problem? Good.

Have a software problem worth solving?

Whether you're starting with an idea, replacing an existing system, or trying to automate an operation that has become too complicated, let's talk.