Extract Data from PDFs with the API2PDF OpenDataLoader API

API2PDF now supports PDF data extraction with three new OpenDataLoader-powered API endpoints.

Developers can provide a URL to a PDF and extract its contents as structured JSON, Markdown, or HTML using a simple REST API.

The three new endpoints are:

  • POST /opendataloader/json — Extract structured JSON from a PDF

  • POST /opendataloader/markdown — Extract Markdown from a PDF

  • POST /opendataloader/html — Extract HTML from a PDF

These endpoints make API2PDF a straightforward PDF data extraction API for developers building AI applications, RAG pipelines, document processing systems, search engines, knowledge bases, data extraction workflows, and other applications that need to turn PDFs into machine-readable content.

Extract PDF to JSON

Use the /opendataloader/json endpoint when you want to extract the contents of a PDF into structured JSON.

Endpoint:

POST https://v2.api2pdf.com/opendataloader/json

Example using cURL:

curl -X POST "https://v2.api2pdf.com/opendataloader/json" \
  -H "Authorization: YOUR-API-KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/document.pdf",
    "fileName": "document.json"
  }'

Replace YOUR-API-KEY with your API2PDF API key and provide the URL of the PDF you want to parse.

The PDF to JSON endpoint is useful when extracted document content needs to be consumed programmatically by another application.

A typical workflow might look like:

PDF → OpenDataLoader → JSON → Application

JSON output is particularly useful for document processing, data pipelines, automated analysis, and applications that need a structured representation of PDF content.

Extract PDF to Markdown

Use the /opendataloader/markdown endpoint to convert the contents of a PDF into Markdown.

Endpoint:

POST https://v2.api2pdf.com/opendataloader/markdown

Example using cURL:

curl -X POST "https://v2.api2pdf.com/opendataloader/markdown" \
  -H "Authorization: YOUR-API-KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/document.pdf",
    "fileName": "document.md"
  }'

Markdown is particularly useful for AI, LLM, and RAG applications because it provides a lightweight text representation while retaining useful document structure.

A typical AI document pipeline could look like:

PDF → OpenDataLoader → Markdown → LLM

For a RAG application:

PDF → Markdown → Chunking → Embeddings → Vector Database → LLM

This makes the PDF to Markdown endpoint useful for preparing existing PDF documents for ChatGPT, Claude, Gemini, AI agents, semantic search, knowledge bases, and other LLM-powered applications.

Extract PDF to HTML

Use the /opendataloader/html endpoint when you want the extracted contents of a PDF represented as HTML.

Endpoint:

POST https://v2.api2pdf.com/opendataloader/html

Example using cURL:

curl -X POST "https://v2.api2pdf.com/opendataloader/html" \
  -H "Authorization: YOUR-API-KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/document.pdf",
    "fileName": "document.html"
  }'

HTML output is useful when extracted PDF content needs to be displayed on the web, transformed using existing HTML tools, indexed, or processed by another application.

A typical workflow might look like:

PDF → OpenDataLoader → HTML → Web Application

One PDF Extraction API, Three Output Formats

The three OpenDataLoader endpoints provide different representations of the same basic operation: extracting usable content from a PDF.

Choose the output format based on what your application needs.

JSON is a good choice for structured data processing and application-to-application workflows.

Markdown is a good choice for AI, LLM, RAG, semantic search, and text-processing workflows.

HTML is a good choice for web applications, content migration, rendering, indexing, and HTML-based processing.

All three endpoints use the same simple request pattern:

{
  "url": "https://example.com/document.pdf",
  "fileName": "output-file"
}

Your application provides the PDF URL and API2PDF handles parsing and extraction.

PDF Data Extraction for AI and LLM Applications

PDFs contain an enormous amount of information, but PDF itself is often an inconvenient input format for modern AI applications.

LLMs work primarily with text and structured representations of content. Before a PDF can be effectively used by an AI application, its contents usually need to be extracted and transformed into a format suitable for downstream processing.

API2PDF's OpenDataLoader endpoints provide that extraction layer.

For example:

PDF → API2PDF → Markdown → Claude

PDF → API2PDF → Markdown → ChatGPT

PDF → API2PDF → JSON → AI Agent

PDF → API2PDF → Markdown → Embeddings → Vector Database → RAG

This makes it possible to build document-aware AI applications without maintaining your own PDF parsing infrastructure.

Build RAG Pipelines from PDFs

Retrieval-Augmented Generation, or RAG, commonly requires converting source documents into text before they can be chunked, embedded, and indexed.

With the PDF to Markdown endpoint, a RAG ingestion pipeline can begin with a simple API request.

A typical architecture might be:

PDF URL

↓

API2PDF OpenDataLoader

↓

Markdown

↓

Text Chunking

↓

Embeddings

↓

Vector Database

↓

LLM

This approach can be used to build AI search systems and knowledge bases from collections of PDF documents.

Build PDF Data Extraction Workflows

OpenDataLoader isn't limited to AI applications.

The endpoints can also be used anywhere an application needs to extract content from PDFs programmatically.

Common use cases include:

  • PDF data extraction

  • PDF parsing

  • PDF to JSON conversion

  • PDF to Markdown conversion

  • PDF to HTML conversion

  • AI document ingestion

  • RAG pipelines

  • LLM preprocessing

  • AI agents

  • Document analysis

  • Knowledge bases

  • Semantic search

  • Enterprise search

  • Document indexing

  • Content migration

  • Automated document processing

  • Search indexing

  • Data pipelines

Instead of operating and scaling your own PDF parsing infrastructure, your application can send a PDF URL to API2PDF and receive the extracted document in the format appropriate for your workflow.

Access PDFs Behind Authentication

All three OpenDataLoader endpoints support extraHTTPHeaders.

That means API2PDF can retrieve a source PDF that requires additional HTTP headers.

For example:

{
  "url": "https://example.com/private-document.pdf",
  "fileName": "document.md",
  "extraHTTPHeaders": {
    "Authorization": "Bearer YOUR-TOKEN"
  }
}

This can be useful when PDFs are stored behind authenticated endpoints or other systems that require HTTP headers when retrieving a file.

OpenDataLoader as a Hosted API

OpenDataLoader provides document parsing and extraction capabilities for turning PDFs into useful structured content.

With API2PDF, developers can access OpenDataLoader through a hosted REST API without deploying and maintaining their own document parsing infrastructure.

The basic workflow is simple:

PDF URL → API2PDF OpenDataLoader API → JSON, Markdown, or HTML

Your application can then store, transform, index, analyze, display, or send the extracted content to another service.

OpenDataLoader API Endpoint Reference

PDF to JSON

Method: POST

Endpoint: https://v2.api2pdf.com/opendataloader/json

Extracts structured JSON from a PDF URL.

PDF to Markdown

Method: POST

Endpoint: https://v2.api2pdf.com/opendataloader/markdown

Extracts Markdown from a PDF URL.

PDF to HTML

Method: POST

Endpoint: https://v2.api2pdf.com/opendataloader/html

Extracts HTML from a PDF URL.

All three endpoints use API2PDF authentication:

Authorization: YOUR-API-KEY

Requests use:

Content-Type: application/json

The primary request fields are:

url — URL of the PDF file to parse.

fileName — Optional filename for the generated JSON, Markdown, or HTML output.

extraHTTPHeaders — Optional HTTP headers API2PDF should include when retrieving the source PDF.

storage — Optional custom storage configuration for storing the generated output outside API2PDF's default storage.

The endpoints also support the outputBinary query parameter.

Getting Started

If you already use API2PDF, you can start extracting data from PDFs immediately with one of the three new OpenDataLoader endpoints:

POST /opendataloader/json

POST /opendataloader/markdown

POST /opendataloader/html

Choose JSON when you need structured data, Markdown when you're building AI and text-processing workflows, or HTML when the extracted content will be used by a web or HTML-based application.

API2PDF now provides APIs for both sides of automated document workflows: generating PDFs from HTML, URLs, and Markdown, and extracting existing PDFs into JSON, Markdown, and HTML.

For complete request specifications and available options, see the API2PDF v2 documentation.