BeakGraph

RDF datasets as self-contained HDF5 files, queried with SPARQL in place.

Get started View on GitHub


BeakGraph is an Apache Jena graph implementation of RDF HDT technology, stored in an HDF5 file and extended to a full RDF dataset. It converts RDF into dictionary-encoded, columnar quad stores: one .h5 file is one self-contained, immutable store, queried with SPARQL directly from the file — locally or over HTTP range requests — without loading the graph into memory.

Highlights

  • Six conversion engines, one format. In-memory engines (-method 0/2/3) for data that fits in RAM; disk-based engines (-method 1/4/5) for inputs larger than memory, up to 10⁹–10¹¹ quads. Every engine writes the same on-disk format.
  • Query in place. A store on a web server or object store is queried through HTTP range requests without downloading it.
  • SPARQL endpoint and LWS storage. -endpoint serves one store as a read-only SPARQL endpoint, or a whole directory of stores as W3C LWS storage.
  • RDF 1.2. Base-direction literals and triple terms are stored and matched term-exactly; the W3C RDF 1.2 and SPARQL 1.2 test suites run in CI against every writer engine.
  • GeoSPARQL. geof:sfIntersects is answered from a Hilbert-curve spatial index, with every candidate verified against the real geometry, so results are exact.
  • An open format. The file format specification is precise enough to write BeakGraph stores without the BeakGraph sources.

Quick start

Requires Java 25 and Maven 3.9. The disk-based engines also need the native HDF5 library — see Requirements.

# Build the self-contained command-line jar
mvn -Pcmdlinejar clean package

# Convert every RDF file under ./data to one .h5 per file under ./out
java -jar target/BeakGraph-<version>.jar -src data/ -dest out/

# Serve a store as a SPARQL endpoint on port 8888
java -jar target/BeakGraph-<version>.jar -endpoint out/example.h5 -port 8888

From Java, open a store and query it with Jena as usual:

try (BeakGraph bg = BG.getBeakGraph(new File("mydata.ttl.h5"))) {
    Dataset ds = bg.getDataset();
    ds.getDefaultModel().write(System.out, "NTRIPLE");
}

Documentation

Page What it covers
Instructions Building, the command-line reference, choosing a conversion engine, very large builds, exporting, verifying, serving, and using BeakGraph from Java.
File format The on-disk format of a BeakGraph store: HDF5 profile, encodings, term order, dictionary, quad indexes and ingest rules.
RDF 1.2 compliance Conformance to RDF 1.1, RDF 1.2 and SPARQL-CDT, and the documented deviations.
Benchmarks The JMH benchmarks for the read path.
Changelog Release notes and the format version history.
Architecture The original design of the on-disk format, slide by slide: file layout, dictionaries, indexes, rank/select, the query path and the writers.

Background

The general idea of BeakGraph is a read-only, searchable, indexed set of binary succinct data structures representing an RDF dataset. It was developed to power Halcyon; see the paper at arXiv:2304.10612.

BeakGraph is licensed under the Apache License 2.0.