BeakGraph
RDF datasets as self-contained HDF5 files, queried with SPARQL in place.
BeakGraph is an Apache Jena graph implementation of RDF HDT technology, stored in an HDF5 file and extended to a full RDF dataset. It converts RDF into dictionary-encoded, columnar quad stores: one .h5 file is one self-contained, immutable store, queried with SPARQL directly from the file — locally or over HTTP range requests — without loading the graph into memory.
Highlights
- Six conversion engines, one format. In-memory engines (
-method 0/2/3) for data that fits in RAM; disk-based engines (-method 1/4/5) for inputs larger than memory, up to 10⁹–10¹¹ quads. Every engine writes the same on-disk format. - Query in place. A store on a web server or object store is queried through HTTP range requests without downloading it.
- SPARQL endpoint and LWS storage.
-endpointserves one store as a read-only SPARQL endpoint, or a whole directory of stores as W3C LWS storage. - RDF 1.2. Base-direction literals and triple terms are stored and matched term-exactly; the W3C RDF 1.2 and SPARQL 1.2 test suites run in CI against every writer engine.
- GeoSPARQL.
geof:sfIntersectsis answered from a Hilbert-curve spatial index, with every candidate verified against the real geometry, so results are exact. - An open format. The file format specification is precise enough to write BeakGraph stores without the BeakGraph sources.
Quick start
Requires Java 25 and Maven 3.9. The disk-based engines also need the native HDF5 library — see Requirements.
# Build the self-contained command-line jar
mvn -Pcmdlinejar clean package
# Convert every RDF file under ./data to one .h5 per file under ./out
java -jar target/BeakGraph-<version>.jar -src data/ -dest out/
# Serve a store as a SPARQL endpoint on port 8888
java -jar target/BeakGraph-<version>.jar -endpoint out/example.h5 -port 8888
From Java, open a store and query it with Jena as usual:
try (BeakGraph bg = BG.getBeakGraph(new File("mydata.ttl.h5"))) {
Dataset ds = bg.getDataset();
ds.getDefaultModel().write(System.out, "NTRIPLE");
}
Documentation
| Page | What it covers |
|---|---|
| Instructions | Building, the command-line reference, choosing a conversion engine, very large builds, exporting, verifying, serving, and using BeakGraph from Java. |
| File format | The on-disk format of a BeakGraph store: HDF5 profile, encodings, term order, dictionary, quad indexes and ingest rules. |
| RDF 1.2 compliance | Conformance to RDF 1.1, RDF 1.2 and SPARQL-CDT, and the documented deviations. |
| Benchmarks | The JMH benchmarks for the read path. |
| Changelog | Release notes and the format version history. |
| Architecture | The original design of the on-disk format, slide by slide: file layout, dictionaries, indexes, rank/select, the query path and the writers. |
Background
The general idea of BeakGraph is a read-only, searchable, indexed set of binary succinct data structures representing an RDF dataset. It was developed to power Halcyon; see the paper at arXiv:2304.10612.
BeakGraph is licensed under the Apache License 2.0.