# docling-graph **Repository Path**: mirrors_ibm/docling-graph ## Basic Information - **Project Name**: docling-graph - **Description**: Transform unstructured documents into validated, rich and queryable knowledge graphs. - **Primary Language**: Unknown - **License**: MIT - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2025-11-18 - **Last Updated**: 2026-09-20 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README
# Docling Graph [](https://docling-project.github.io/docling-graph/) [](https://pypi.org/project/docling-graph/) [](https://www.python.org/downloads/) [](https://github.com/astral-sh/uv) [](https://github.com/astral-sh/ruff) [](https://opensource.org/licenses/MIT) [](https://pydantic.dev) [](https://github.com/docling-project/docling) [](https://networkx.org/) [](https://typer.tiangolo.com/) [](https://github.com/Textualize/rich) [](https://vllm.ai/) [](https://ollama.ai/) [](https://www.bestpractices.dev/projects/11598) [](https://lfaidata.foundation/projects/) Docling-Graph turns documents into validated **Pydantic** objects, then builds a **directed knowledge graph** with explicit semantic relationships. This transformation enables high-precision use cases in **chemistry, finance, and legal** domains, where AI must capture exact entity connections (compounds and reactions, instruments and dependencies, properties and measurements) **rather than rely on approximate text embeddings**. This toolkit supports two extraction paths: **local VLM extraction** via Docling, and **LLM-based extraction** routed through **LiteLLM** for local runtimes (vLLM, Ollama) and API providers (OpenAI, Gemini, IBM watsonx, Mistral and more), all orchestrated through a flexible, config-driven pipeline. ## Key Capabilities - **✍🏻 Input formats:** [Docling](https://docling-project.github.io/docling/usage/supported_formats/)’s supported inputs: PDF, images, DocLang, markdown, Office and more. - **🧠 Extraction:** [LLM](https://docling-project.github.io/docling-graph/fundamentals/pipeline-configuration/backend-selection/) or [VLM](https://docling-project.github.io/docling-graph/fundamentals/pipeline-configuration/backend-selection/) backends, with [chunking](https://docling-project.github.io/docling-graph/fundamentals/extraction-process/chunking-strategies/) and [processing modes](https://docling-project.github.io/docling-graph/fundamentals/pipeline-configuration/processing-modes/). - **💎 Graphs:** Pydantic to [NetworkX](https://docling-project.github.io/docling-graph/fundamentals/graph-management/graph-conversion/) directed graphs with stable IDs, edge and [provenance](https://docling-project.github.io/docling-graph/fundamentals/graph-management/provenance/) metadata. - **📦 Export:** [CSV](https://docling-project.github.io/docling-graph/fundamentals/graph-management/export-formats/#csv-export), [Cypher](https://docling-project.github.io/docling-graph/fundamentals/graph-management/export-formats/#cypher-export), and other KG-friendly formats. - **🔍 Visualization:** [Interactive HTML](https://docling-project.github.io/docling-graph/fundamentals/graph-management/visualization/) and Markdown reports. - **🐛 Trace capture:** [Debug exports](https://docling-project.github.io/docling-graph/usage/advanced/trace-data-debugging/) for extraction and fallback diagnostics. ### Latest Changes - **🔗 Graph fusion:** [Merge](https://docling-project.github.io/docling-graph/usage/cli/merge-command/) multiple knowledge graphs into one. Fully audited, deterministic, and no LLM calls. - **🧩 Template generation:** [Generate](https://docling-project.github.io/docling-graph/usage/cli/template-command/) Pydantic templates from example documents or ontologies (OWL/RDFS...). - **🦆 DocLang support:** Parse `.dclg`/`.dclx` inputs, and [optionally serialize](https://docling-project.github.io/docling-graph/fundamentals/extraction-process/document-conversion/#llm-input-serialization) document as [DocLang](https://github.com/doclang-project/doclang) for the LLM. - **📍 Data grounding:** Deterministic [provenance](https://docling-project.github.io/docling-graph/fundamentals/graph-management/provenance/) ledger with bounding-box geometry and no extra LLM calls. - **✨ Dense extraction:** Advanced [skeleton-then-flesh](https://docling-project.github.io/docling-graph/fundamentals/extraction-process/dense-extraction/) extraction mode for complex documents. - **🚀 Docling Serve support:** Offload [document conversion](https://docling-project.github.io/docling-graph/fundamentals/pipeline-configuration/docling-serve/) to a remote [docling-serve](https://github.com/docling-project/docling-serve) instance. ## Quick Start ### Requirements - Python 3.10 or higher ### Installation ```bash pip install docling-graph ``` This installs the core package with LiteLLM for remote and local LLM providers. VLM backend support requires the `vlm` extra: ```bash pip install "docling-graph[vlm] ``` For detailed installation instructions (including optional extras and GPU setup), see [Installation Guide](https://docling-project.github.io/docling-graph/fundamentals/installation/). ### API Key Setup (Remote Inference) Copy [`.env.example`](.env.example) to `.env` and fill in the values for the provider(s) you use: ```bash cp .env.example .env ``` See [API Keys Setup](https://docling-project.github.io/docling-graph/fundamentals/installation/api-keys/) for provider-specific instructions (including Amazon Bedrock's AWS credential chain). ### Basic Usage #### CLI ```bash # Initialize configuration docling-graph init # Convert document from URL (each line except the last must end with \) docling-graph convert "https://arxiv.org/pdf/2207.02720" \ --template "docs.examples.templates.rheology_research.ScholarlyRheologyPaper" \ --processing-mode "many-to-one" \ --extraction-contract "dense" \ --debug # Visualize results docling-graph inspect outputs ``` #### Python API - Default Behavior ```python from docling_graph import run_pipeline, PipelineContext from docs.examples.templates.rheology_research import ScholarlyRheologyPaper # Create configuration config = { "source": "https://arxiv.org/pdf/2207.02720", "template": ScholarlyRheologyPaper, "backend": "llm", "inference": "remote", "processing_mode": "many-to-one", "extraction_contract": "auto", "provider_override": "mistral", "model_override": "mistral-medium-latest", "structured_output": True, # default "use_chunking": True, } # Run pipeline - returns data directly, no files written to disk context: PipelineContext = run_pipeline(config) # Access results graph = context.knowledge_graph models = context.extracted_models metadata = context.graph_metadata print(f"Extracted {len(models)} model(s)") print(f"Graph: {graph.number_of_nodes()} nodes, {graph.number_of_edges()} edges") ``` Every node above also carries a deterministic `__provenance__` attribute by default (`provenance="standard"`), pointing back to the source chunk and page it was extracted from — no extra LLM calls involved. See [Data Grounding & Provenance](https://docling-project.github.io/docling-graph/fundamentals/graph-management/provenance/). For debugging, use `--debug` with the CLI to save intermediate artifacts to disk; see [Trace Data & Debugging](https://docling-project.github.io/docling-graph/usage/advanced/trace-data-debugging/). For more examples, see [Examples](https://docling-project.github.io/docling-graph/usage/examples/). ## Pydantic Templates Templates define both the **extraction schema** and the resulting **graph structure**. ```python from pydantic import BaseModel, Field from docling_graph.utils import edge class Person(BaseModel): """Person entity with stable ID.""" model_config = { 'is_entity': True, 'graph_id_fields': ['last_name', 'date_of_birth'] } first_name: str = Field(description="Person's first name") last_name: str = Field(description="Person's last name") date_of_birth: str = Field(description="Date of birth (YYYY-MM-DD)") class Organization(BaseModel): """Organization entity.""" model_config = {'is_entity': True} name: str = Field(description="Organization name") employees: list[Person] = edge("EMPLOYS", description="List of employees") ``` ### Generating a template from documents Instead of writing the template by hand, you can induce one from a few example documents: ```bash docling-graph template from-docs invoice1.pdf invoice2.pdf \ --output templates/invoices.py \ --name InvoiceDocument \ --trial-run ``` The documents are converted with Docling, then LLM passes propose classes, fields, and relationships **as structured data** — a deterministic renderer turns that into the Python module, so no LLM ever writes code. Candidates are filtered by deterministic gates (every identity example must appear verbatim in the source) and merged across documents. `--trial-run` then runs a real extraction on the first document and prints an advisory quality report. Each generator also writes an editable SPEC YAML next to the template (`templates/invoices.spec.yaml`). Rename an edge or flip an entity to a component with a one-line YAML edit and re-render, rather than hand-patching generated code: ```bash docling-graph template from-spec templates/invoices.spec.yaml -o templates/invoices.py ``` Templates can also be compiled from an existing ontology — OWL/RDFS/SKOS, LinkML, or JSON Schema — with no LLM involved at all (needs the `templategen` extra: `pip install 'docling-graph[templategen]'`). Any template, generated or hand-written, can be checked against the rulebook: ```bash docling-graph template from-ontology schema.ttl --root ex:InsurancePolicy -o templates/policy.py docling-graph template lint templates.invoices.InvoiceDocument ``` For complete guidance, see: - [template Command](https://docling-project.github.io/docling-graph/usage/cli/template-command/) — generating, linting, and evaluating templates - [Schema Definition Guide](https://docling-project.github.io/docling-graph/fundamentals/schema-definition/) - [Template Basics](https://docling-project.github.io/docling-graph/fundamentals/schema-definition/template-basics/) - [Example Templates](docs/examples/README.md) ## Documentation Comprehensive documentation can be found on [Docling Graph's Page](https://docling-project.github.io/docling-graph/). ### Documentation Structure The documentation follows the docling-graph pipeline stages: 1. [Introduction](https://docling-project.github.io/docling-graph/introduction/) - Overview and core concepts 2. [Installation](https://docling-project.github.io/docling-graph/fundamentals/installation/) - Setup and environment configuration 3. [Schema Definition](https://docling-project.github.io/docling-graph/fundamentals/schema-definition/) - Creating Pydantic templates 4. [Pipeline Configuration](https://docling-project.github.io/docling-graph/fundamentals/pipeline-configuration/) - Configuring the extraction pipeline 5. [Extraction Process](https://docling-project.github.io/docling-graph/fundamentals/extraction-process/) - Document conversion and extraction 6. [Graph Management](https://docling-project.github.io/docling-graph/fundamentals/graph-management/) - Converting, grounding, exporting, and visualizing graphs 7. [CLI Reference](https://docling-project.github.io/docling-graph/usage/cli/) - Command-line interface guide 8. [Python API](https://docling-project.github.io/docling-graph/usage/api/) - Programmatic usage 9. [Examples](https://docling-project.github.io/docling-graph/usage/examples/) - Working code examples 10. [Advanced Topics](https://docling-project.github.io/docling-graph/usage/advanced/) - Performance, testing, error handling 11. [API Reference](https://docling-project.github.io/docling-graph/reference/) - Detailed API documentation 12. [Community](https://docling-project.github.io/docling-graph/community/) - Contributing and development guide ## Contributing We welcome contributions! Please see: - [Contributing Guidelines](.github/CONTRIBUTING.md) - How to contribute - [Development Guide](https://docling-project.github.io/docling-graph/community/) - Development setup ### Development Setup ```bash # Clone and setup git clone https://github.com/docling-project/docling-graph cd docling-graph # Install with dev dependencies uv sync --extra dev # Run Execute pre-commit checks uv run pre-commit run --all-files ``` ## License MIT License - see [LICENSE](LICENSE) for details. ## Acknowledgments Docling Graph builds on outstanding open-source projects: - [Docling](https://github.com/docling-project/docling) - document conversion and VLM extraction - [Pydantic](https://pydantic.dev) - schema definition and validation - [NetworkX](https://networkx.org/) - graph construction and analysis - [LiteLLM](https://github.com/BerriAI/litellm) - unified LLM provider interface - [Cytoscape](https://js.cytoscape.org/) - interactive graph visualization ## IBM ❤️ Open Source AI Docling Graph has been brought to you by IBM.