Skip to main content
ACM Platforms

Coming soon

Document Parser for RAG

Enterprise document-to-RAG pipeline — turn textbooks, technical documents, and scanned PDFs into structured, chunked, knowledge-base-ready content with math-preserving fidelity and deterministic reproducibility.

The problem we solve

Without Document Parser

  • Monolithic PDFs fed into knowledge bases lose section context
  • Math formulas and tables stripped during chunking
  • Answer keys exposed to student-facing AI tutors
  • Re-ingesting one corrected page means re-processing the entire book
  • No metadata for chapter/topic-level retrieval filtering

With Document Parser for RAG

  • Per-section files aligned to hierarchical chunking boundaries
  • LaTeX math and tables preserved through the entire pipeline
  • Answer keys isolated in a separate prefix with access control
  • Granular re-ingestion — one section, one file, one update
  • Rich metadata sidecars for chapter/topic/content-type filtering

Independently validated conversion quality

Three OpenStax STEM textbooks converted and evaluated by an independent AI reviewer against the original PDFs — scoring OCR accuracy, markdown structure, math notation, tables, figures, educational structure, and RAG readiness.

OpenStax Physics

875 pages

9.95/10

Overall fidelity score

RAG Readiness9.98/10

Highest-quality conversion

OpenStax Intermediate Algebra 2e

1,335 pages

9.93/10

Overall fidelity score

RAG Readiness9.95/10

Production-quality

OpenStax Calculus Volume 1

769 pages

9.7/10

Overall fidelity score

RAG Readiness9.8/10

Very high-quality

What the evaluations measure

OCR Accuracy

Text fidelity vs. original PDF

Math Notation

LaTeX preservation for formulas

Document Structure

Heading hierarchy and reading order

RAG Readiness

Chunk boundaries, embedding quality, metadata

Evaluated across 2,979 pages of OpenStax STEM content including hundreds of diagrams, complex tables, multi-line derivations, Greek symbols, scientific notation, and hundreds of worked examples. Reports conclude the pipeline can achieve 99–99.5% fidelity with targeted post-processing refinements.

Supported formats

Ingest from six source formats; export to eight destination formats — each selectable per job via the API or portal UI.

Input formats

  • PDFDigital, scanned, and mixed — with layout-aware OCR fallback
  • DOCXMicrosoft Word with tables, headings, and images
  • EPUBE-books with full structural hierarchy
  • MarkdownPass-through with structure validation
  • HTMLWeb pages and exported documents
  • TXTPlain text with heuristic sectioning

Output formats

  • MarkdownMath-preserving LaTeX ($…$, $$…$$) and tables intact
  • Bedrock KBPer-section .md + metadata sidecar, hierarchical chunking ready
  • JSONLOne JSON object per section — RAG-ready chunks
  • JSONFull structured representation of the parsed document
  • ParquetColumnar format for analytics and vector pipelines
  • HTML / YAML / XML / CSVAdditional export targets for downstream systems

Core capabilities

Built around one principle: deterministic parsing is separated from AI enrichment. Structure extraction is reproducible and verifiable; AI enrichment is optional and independently triggered.

Deterministic parsing

Structure extraction, validation, and format conversion are reproducible and hallucination-free. Same input + same settings → same output every time.

AI enrichment (opt-in)

Summaries, flashcards, quizzes, and learning objectives generated on demand — pluggable between a local model and AWS Bedrock. Never mixed into the conversion path.

Math & formula fidelity

LaTeX expressions preserved end-to-end through the pipeline. Bedrock KB exporter retains $…$ inline and $$…$$ display math — critical for STEM RAG.

Content-type classification

Lessons, examples, exercises, quizzes, and answer keys automatically classified — enabling tutor agents to filter retrieval and keep answer keys out of student-facing responses.

Metadata sidecars

Each output file paired with structured metadata for AWS Bedrock Knowledge Base filtering — book, chapter, section, page range, and content type available at query time.

Multi-tenant SaaS

Organization isolation, role-based access, API keys, subscription tiers, and per-tenant storage — ready for white-label and enterprise deployment.

Purpose-built for RAG pipelines

Designed specifically for AWS Bedrock Knowledge Bases — output structure, metadata, and content-type classification optimized for retrieval-augmented generation with AI tutoring agents.

Hierarchical chunking alignment

Output structured so AWS Bedrock Knowledge Base hierarchical chunking lands on natural section boundaries — maximizing retrieval precision for AI agents.

Answer-key isolation

Answer keys routed to a separate, access-controlled location — configurable as a restricted data source or excluded entirely — preventing AI tutors from handing students raw answers.

Metadata-filtered retrieval

Per-file attributes (book, chapter, section, content type, page range) enable agents to scope queries to specific chapters or content types at runtime.

Targeted re-ingestion

Editing one section re-ingests one file — not the entire book. Granular updates without full re-vectorization.

How it works

From raw document to knowledge-base-ready output — each stage is observable, configurable, and delivers verifiable results.

01

Ingestion & routing

  • Automatic source format detection (PDF, DOCX, EPUB, Markdown, HTML, TXT)
  • Document profile selection: textbook, receipt, general scan, born-digital
  • Quality tier routing: fast vs. high-fidelity for your throughput and accuracy needs
02

Parsing & extraction

  • Layout-aware OCR for scanned pages with math and table recognition
  • Digital PDF extraction preserving structure, headings, and figures inline
  • Fallback pipelines for degraded scans with image preprocessing
  • Native support for Word, EPUB, Markdown, HTML, and plain text
03

Structured representation

  • Unified document model: chapters → sections → blocks (paragraph, heading, table, formula, image, code)
  • Block-level metadata: page numbers, positions, and confidence scores
  • Automated validation report with quality scoring and actionable issue list
04

Export & delivery

  • Format-selective export: choose one or many destination formats per job
  • AWS Bedrock Knowledge Base output: per-section files + metadata sidecars + manifest
  • Secure cloud storage with presigned download URLs and async job processing

Built on AWS

Production-grade infrastructure on Amazon Web Services — the same enterprise cloud foundation as our live platforms.

AI & ML

AWS Bedrock for AI enrichment and Knowledge Base integration. Local model option available for air-gapped deployments.

Secure storage

Encrypted document storage with presigned download URLs and configurable retention policies.

Async processing

Queue-based job architecture with automatic retries, dead-letter handling, and horizontal scaling.

Enterprise auth

AWS Cognito integration with role-based access, multi-tenant isolation, and API key access for automation.

Use cases

Education & tutoring

Convert textbooks into knowledge bases for AI tutoring agents — with answer-key isolation ensuring students learn through guided discovery, not copied answers.

Enterprise documentation

Turn technical manuals, compliance documents, and internal wikis into searchable AI knowledge bases with chapter-level filtering and access control.

Research & publishing

Process academic papers and research libraries into structured corpora — preserving formulas, citations, and figure references for domain-specific RAG.

Deployment options

Same ACMP pattern as our live platforms — SaaS, self-hosted, and managed dedicated tiers.

SaaS (multi-tenant)

Self-serve

Upload via portal or API. Per-tenant isolation, subscription tiers, and usage-based billing.

Self-hosted

Your infrastructure

Deploy in your VPC with your own cloud resources. You own compute and data residency.

Managed dedicated

ACM-managed

Isolated environment managed by ACM Platforms. SLA-backed, usage-based pricing, scaling and patching included.

Ready to turn your documents into AI-ready knowledge?

Join our early access program to be among the first to use the Document Parser for RAG platform. Discuss your use case with our solutions team — education, enterprise, or research.