Skip to content
#

unstructured-io

Here are 17 public repositories matching this topic...

Extract data from images, pdf, invoices, receipts | Extract tables from pdf, images and convert to Excel/CSV | OCR complex pdfs, images.

  • Updated Jan 28, 2026

Developed an end-to-end Multimodal Retrieval-Augmented Generation (RAG) ingestion pipeline for processing complex PDF documents containing text, tables, and images. • Implemented high-resolution PDF parsing using Unstructured.io with OCR support for scanned documents, enabling extraction text, tables, and embedded images while preserving document

  • Updated Aug 5, 2026
  • Jupyter Notebook

Documentation assistant for developers who want to quickly understand and query large documentation sites. Built with a modern tech stack including Firecrawl for llm-ready web crawling, Unstructured for document processing, MongoDB Atlas for vector search, and OpenAI for embeddings and generation.

  • Updated Jan 16, 2026
  • Python

Add this topic to your repo

To associate your repository with the unstructured-io topic, visit your repo's landing page and select "manage topics."

Learn more