Most document-extraction stacks make you choose: cloud APIs that see your data, or brittle local scripts. xberg is a self-hosted ingestion engine for AI systems. It turns private documents and websites into structured, citation-backed, LLM-ready data from one local container. Rust core (pdf_oxide, SIMD, full parallelism), 96 file formats, 16 language SDKs from one engine. Local-first and CPU-only: thousands of documents a minute, no GPU, and your data never leaves your infrastructure. Docs in the first comment.
Xberg.io
Softwareentwicklung
Self-hosted ingestion engine for AI systems: documents and websites to structured, LLM-ready data
Info
Xberg is the self-hosted ingestion engine for AI systems. It turns private documents and websites into structured, citation-backed, LLM-ready data from one local container. Rust core, 96 file formats, 16 language SDKs, local-first and CPU-only, so your data never leaves your infrastructure. Built on a simple idea: getting clean, structured data out of documents and the web shouldn't force a choice between cloud APIs that see your data and fragile local scripts.
- Website
-
https://docs.xberg.io
Externer Link zu Xberg.io
- Branche
- Softwareentwicklung
- Größe
- 2–10 Beschäftigte
- Hauptsitz
- Berlin
- Art
- Privatunternehmen
- Spezialgebiete
- Data, Data Processing, Data Intelligence, RAG, AI, PDF Parsing, RAG Infrastructure, OCR Processing, API, Open Source, Data Extraction, Preprocessing Pipeline, Data Pipeline, RAG Pipeline, Rust, Python, Machine Learning, Artificial Intelligence, PDF Extraction und Metadata Extraction
Orte
-
Primär
Wegbeschreibung
Berlin, DE
Beschäftigte von Xberg.io
Updates
-
We’ve released html-to-markdown v3.8.0. This release focuses on developer experience and packaging: • Consolidated Node package distribution • Corrected PHP install path for native extension loading (PIE) • No conversion behavior/output changes Read more: https://lnkd.in/dr6s7QjC #OpenSource #RustLang #NodeJS #PHP #DeveloperTools #WebDev
-
-
Time to make it 10k; This weekend. Let's go. ⭐ https://lnkd.in/ds_r9Pss
-
-
What happens to my documents? 🐙 Documents are processed in memory and deleted immediately after extraction. No storage, no indexing. We don't train on your data or use it for model improvement. kreuzberg.dev #documentintelligence #AIinfrastructure #RAG #agenticAI #documentprocessing
-
Does Kreuzberg handle scanned documents? 🐙 Yes. Built-in OCR recognizes text in images and scanned PDFs. No additional configuration is needed. Just send the file and get structured output back. kreuzberg.dev #documentintelligence #AIinfrastructure #RAG #agenticAI #documentprocessing
-
How fast is 'fast'? Kreuzberg is built on a high-performance Rust core, so most documents are processed almost instantly- in milliseconds instead of seconds. For bulk jobs that's thousands of pages per hour on a single API key. Benchmarks coming up soon. Kreuzberg.dev #documentintelligence #AIinfrastructure #RAG #agenticAI
-
FYI: You can get Kreuzberg as a docker! The open source is widely available! We have official docker files, a CLI tool (you can install it with brew), and bindings for 16 programming languages. https://kreuzberg.dev/ https://lnkd.in/dAADu9PT
I was asked if you can get Kreuzberg as docker. Yes! The open source is widely available! We have official docker files, a CLI tool (you can install it with brew), and bindings for 16 programming languages. see more at https://lnkd.in/d6_JvD4t
-
We built the fastest document extraction engine available. Then we made it a managed API. Kreuzberg Cloud is live. 97+ formats, sub-second latency, structured output ready for any pipeline or agent. Built on the OSS library that's already running in production at scale. First 10K pages free. No card. kreuzberg.dev #documentintelligence #AIinfrastructure #RAG #agenticAI