Changelog
Notable changes to BeakGraph and its on-disk format. The format history is normative (SPECIFICATIONS.md §4.1); the rest is a summary of what a release brought and what a reader must rebuild.
Unreleased (branch rdf12andcdt, after 0.18.0)
- Jena 6.2.0, jHDF 0.13.0.
- LWS storage mode follows the LWS Protocol editor’s draft of 2026-10-05 for everything a read-only storage serves. Every GET/HEAD response links the storage URI (
/description) withrel="https://www.w3.org/ns/lws#storage"(was…#storageDescription); the storage URI serves the storage description asapplication/lws+cid, a Controlled Identifier document (@context[cid/v1, lws/v1],idthe storage URI) with the mandatoryStorageRootservice beside the SPARQL endpoint. Container items carryformat(wasmediaType) and the JSON body no longer mirrors pagination asas:first/as:next/… (Link headers only);application/ld+jsonwith the LWS profile is echoed. Linksets hold only link relations (type, up, linkset, storage) and, like the description, carry ETag and Last-Modified. PUT, PATCH, DELETE, TRACE and POST (except a SPARQL query to a store) answer 405 withAllow; OPTIONS reports it; errors are RFC 9457 problem details; If-Match and If-Unmodified-Since are evaluated (412); a container’s HTML, Turtle and JSON listings have distinct ETags. Single-file mode no longer serves a storage description - it has no storage. The metadata model (beakgraph.ttl.gzand the directory-mode/rdfdataset) uses the LWS vocabulary’sdcterms:format,dcterms:modifiedandlws:totalItemsinstead ofas:mediaType,as:updatedandas:totalItems: a cache written with the old terms is regenerated at start-up, and SPARQL queries over/rdfnaming them must be updated. - Vendored zstd: the pure-Java compressor keeps one compression context per window size and reuses it across frames (hash/chain tables, sequence store and entropy workspaces used to be allocated for every compressed dictionary fragment); frames are byte-identical to before, and a compressor instance is now explicitly single-threaded (one per writer, one per reader thread, as before).
- Housekeeping: the bit-packed buffer’s word fetch dispatches on the sealed
RandomAccessBytesimplementation (a JVM that had opened small, FFM-mapped and remote datasets paid a megamorphic call on every id lookup: +57% on a random get, +42% on a binary search in the benchmark); the buffer’s write-side constructor is(Path, int bitWidth)and its cursor reads,HDTBitmapDirectory.getIds(),DataOutputBuffer.writeInt,NodeSorter.sort2and UTIL’s binary-dump helpers are gone; the in-memory buffers declare no checked exceptions on construction or close;Statsreports “-“ and the allocated width instead of its seed sentinels; the native HDF5 backend names the full path in a failed create; dead code and stale javadoc removed (BGVoIDSD, OpExecutorBG). - Alternative writer engines: one publish discipline (AtomicPublish.build) for all six engines, one spill/merge scaffolding for the two -method 4 sorters, one index level-emission rule for methods 0/1/2 and one id-space helper for the three dictionary writers; the packed sorters reject an id wider than its declared width; the ultra index pads with word-wise range fills; a remote JSON-LD context is fetched once per process, shared and coalesced across concurrent parses; the ultra ingest no longer shadows the base builder’s configuration and refuses build(); the disk builders share a base with documented defaults; docs state what runs outside
-cores. - Default writer engine: the GSPO/GPOS indexes sort and scan dictionary ids resolved once per quad instead of re-comparing nodes and re-searching the dictionaries (same bytes, much less work); the dictionary build refuses two entries the comparator cannot separate; a spatial literal’s later parts keep their index cells when an earlier part’s pyramid fails; a missing language tag fails the build instead of storing “no tag”; the VoID
sd:DatasetIRI is configurable (-voidbase,setVoidDatasetIri) and defaults tourn:x-beakgraph:datasetinstead of the author’s domain; a writer snapshots its builder and an ingest builder is single-use; parse and guard failures keep their own message instead of “I/O error while reading RDF source”; exports use a unique temp name; predicates are guarded at ingest. - Core library and utils: a zip source’s document is chosen by name over the whole archive (
__MACOSX/and dot-file entries skipped, an RDF-named entry preferred, several candidates rejected) instead of taking the first file entry; source streams are buffered and a rejected.gzreleases its file handle; an empty polygon member no longer indexes as Hilbert cell 0; VByte rejects sequences over nine bytes; a corrupt zstd frame surfaces as oneIllegalArgumentExceptionwhichever codec (native or Java) decoded it, and the docs state that compressed bytes depend on the codec (io.airlift.compress.v3.disable-native=trueforces the Java one). - The reference comparator is a strict total order: numbers of every XSD numeric datatype order by exact value (range pushdown widens its bounds to ARQ’s promoted comparison), dateTime and the g* kinds share one instant space with a fixed kind rank, and cross-store bindings are re-resolved. Stores built by earlier 0.18.0 builds that mix numeric datatypes or g* and dateTime literals under one predicate should be rebuilt.
- Relative IRIs resolve against a deep sentinel base, so
<../x>and</x>keep their parent / path-absolute forms;-mergestores each document’s references relative to-src. Earlier stores collapsed them onto the child form (SPECIFICATIONS.md §9.2). - Composite (cdt:) literal constants no longer drive range pushdown (their dictionary order is lexical, ARQ’s is by value).
- Disk-based writers: term sorters spill on a byte budget as well as a record count, and the spill sizing is on the command line (
-spillMB,-termSpillBatch,-idSpillBatch,-mergeFanIn); JSON-LD sources, which Jena parses whole in memory, are warned about and parsed one at a time under-method 5; every engine sorts literals through a per-sort memoizing comparator; the native HDF5 backend writes attributes through thebyte[]entry points, so the-Dhdf5.ffm=truebinding works. -verifyrequires both indexes of a non-empty store, checks every index level’s datasets and declared sizes, and probes dictionary order and searchability (all of it under-deep, which now also re-derives every graph’s rows through GPOS); the-srcwalk skips unreadable subtrees and junction cycles instead of aborting the run.-method 4/5: a failed build drains its spill, merge and parse workers before removing the workspace (no stranded.bghugeultra-*/.bgplaid-*directories); merge levels run at mostmin(cores, 1024 / (fanIn + 1))groups at a time (setMergeConcurrency); the HDF5 backend is probed before parsing.-method 3releases its packed key arrays and rank maps before writing the file.- Disk-based writers: a failed stage cancels its sibling stages at once; a spill run or record file cut short is reported instead of read as a shorter run; every workspace file is closed on failure; the exact output path is probed before parsing; constructors release what they opened. All engines share one relativizer, one literal-stats routing, one spatial augmenter and one dictionary node encoder. Every builder’s
build()names a missing source or destination. - CLI: a bare or mode-less invocation fails with usage (exit 1) instead of exiting 0 silently;
-portis validated;-mergetreats a trailing separator or a suffix-less-destas a directory;-exportand the LWS servlet accept.hdf5like-verify;-export JSON-LDrefuses triple-term stores;-forcerebuilds existing destinations and skips no longer count as successes;-verifyreports a lost or unreadableformatVersionand, under-deep, rows exceedingnumQuads. - The endpoints honour the SPARQL Protocol
default-graph-uri/named-graph-uriparameters;-export -base;-verifyof degenerate stores; the empty-store export fast path; locale-independent identifiers and case mapping throughout. - Query engine:
FROM/FROM NAMED(and the protocol dataset parameters) now run on BeakGraph views - a graph-set view for severalFROMgraphs - instead of Jena’sGraphUnionRead, keeping id-level joins, pushdown and the fast paths;?s ?p <o>is answered by per-predicate GPOS probes instead of a graph scan; range and spatial-index hints survive filter placement after join reordering (they were dropped whenever the placed filter ended up inside an OpSequence). SPECIFICATIONS.md §8.5 gains the access-path table. A parallel scan whose consumer is dropped unclosed is stopped once it is garbage-collected (its workers used to pin it), and scans are planned only for a top-level execution - not per outer row of an OPTIONAL / EXISTS /GRAPH ?gsub-pattern - and never for a store read through an HTTP channel; the spatial-index candidate collection runs at the first row and honours query timeout / abort; the ARQ-global stage generator (general datasets holding BeakGraph views) reorders BGPs like the BG executor; a dictionary read failure while snapping a range filter’s bound now fails the query instead of silently narrowing the range. Range pushdown skips mixed-duration constants ("P1M35D"), whose XSD order the months-first dictionary order can reverse. FFM-mapped datasets (over thebeakgraph.ffm.threshold) are mapped through jHDF’s open channel, so a store rebuilt in place while a reader is open stays consistently on the old file instead of decoding old metadata against new bytes. - Query engine: FILTER range hints are resolved once per store and pattern shape (
RangeBounds, signed with clamped floors) and shared by every iterator, join input row and parallel-scan chunk, replacing four drifted copies; the object-side iterator binary-searches a subject hint; a chunk of a closed scan skips its setup; the first-level and child block-range rules live in one place; the export’s text memo clamps an oversized cache setting and indexes predicate text by id; the node table resolves a miss with one cache operation; a polygon outside the Hilbert domain keeps its clamped cover; the thirdgeof:sfIntersectsargument is refused instead of ignored; row bindings compute their own variables once. - Server: the LWS servlet’s public base and storage root are per instance;
Acceptis negotiated by media range and quality (a plainapplication/jsonrequest is labelled as such);If-None-Matchaccepts lists, weak validators and*; data responses open the file before deriving the validator, are served from that handle, carryCache-Control: no-cacheand an RFC 6266/8187Content-Disposition; a stored entry named*.metais reachable and only the exactHalcyonStoragesegment is an alias; container listings are memoised per metadata snapshot and report live file sizes and times; hidden entries and root entries named like a fixed route are not indexed;void:entitiescounts a multi-typed entity once; the user-profile JSON-LD frame is applied by the response filter instead of a JVM-global writer hook, never fetches a context and never closes the response stream; HDF5 files are typedapplication/x-hdf5and the common RDF, text and image extensions get their media types. INSTRUCTIONS gains “Deployment and trust model”. - Writers: JSON-LD
@contextreferences load from the source tree (a relative reference next to the document, or below-srcfor-merge, used to resolve against the sentinel base and fail); remote contexts are refused unless-jsonLdRemoteis given, and then fetched with a timeout; afile:context must stay inside the tree. A failed ingest closes the parser stream, so Jena’s parser thread no longer stays parked on its queue. Document-relative datatype IRIs ("7"^^<scoreType>) are stored relative and served resolved instead of leaking the sentinel host; a relative IRI inside a composite (cdt:) literal is rejected at ingest like a blank node. Spatial indexing covers every leaf of a nested GEOMETRYCOLLECTION (a nested MULTIPOLYGON was unfindable), and the tile pyramid skips levels over 65,536 tiles instead of enumerating a wide geometry for hours. - Comparator: a literal whose value cannot be built is classified as unparseable against every partner and logged once per datatype; a comparison that fails is an error. The former catch-all ordered that one pair by term - the cyclic value/term mix the CDT and language branches avoid. The cross-space rank table (SPECIFICATIONS.md §6.2) and the
DataTypeordinal table (§7.4) are pinned by tests. - Endpoints:
-timeoutnow also bounds directory mode’s/rdfmetadata dataset; CONSTRUCT / DESCRIBE stream as Turtle and N-Triples and are capped for JSON-LD and RDF/XML (beakgraph.query.construct.max.triples); ASK answers carry the mandatoryheadmember and honourAccept(XML by default, like SELECT); a query may name the served document by relative reference (<>,<image.png>); the storage description advertises/rdfas the SPARQL service and/sparqlredirects to/sparql/; LWS ids, Link targets and RDF subjects are percent-encoded; the JSON-LD user-profile frame applies to SPARQL responses only; single-file mode no longer loads a parent directory’s metadata cache; an unreadable cache is regenerated; an unreadable entry no longer aborts the metadata scan;void:vocabularyis derived from predicate and class namespaces (object paths were unbounded). - Core API: a closed BeakGraph reports
isClosed()and refuses reads, and a closed reader fails every storage call withClosedExceptioninstead of serving stale buffers or a jHDF error; writes are denied with Jena’sAddDeniedException/DeleteDeniedExceptionin both call forms; a named-graph view’s dataset answers the view’s graph as its default graph infind,findNGandlistGraphNodes; a row whose ids cannot be resolved is dropped by the dataset graph as it is by the graph; the constructor’s document base is honoured - the query engine rewrites absolute IRIs under it to the stored relative form and resolves every result row, and the dataset graph andfind()do the same - with a remote (http) store resolving against its own URL by default;RelativeIRIResolvermoved tocore(thecore.fusekiclass is a deprecated alias); the stage-generator director is re-installed when another component replaces the global slot and rides in every BeakGraph dataset context;setDataset(Dataset)is gone from the writer builders (inputs are files); the writers no longer implement the read-side dictionary interfaces (theextract*methods are gone,DictionaryWritergainedlocate); VoID reorder statistics use the distinct-subject and distinct-object role lists; the reader pool accepts http(s) keys and names the schemes it accepts; the HTTP channel follows redirects once at open and re-resolves an expired target, requires and fully checks the Content-Range (and the Content-Length) of a 206, treats a zero Content-Length on HEAD as unknown, retries with exponential backoff and the server’s Retry-After (up to five attempts), and counts every range request it issues. - Readers: one profile helper opens every dataset and attribute, so a chunked dataset or a 32-bit
numEntriesis reported by name instead of a cast error, and integer attributes of any width are accepted; a file with.BGbut no dictionary, an index level whose bitmap and id list differ in length, an FCD group whose block count, offsets or flags do not match its entries, or a triple-term store that disagrees with the datatypes column fails the open; language tags are decoded once at open and must be in Jena’s formatted form; a pre-v3 store logs the linear-select fallback it runs under. The union / any-graph fan-outs drive the per-graph reads from the columnar graph ids, the object dictionary of an all-IRI store answers the true insertion point for literal probes (range hints empty the scan instead of walking it),isGraphis a binary search over the stored list, a literal search reuses the tier terms’ values and builds numeric probe values from the packed numbers, FCD blocks decode without per-entry copies, bit-packed and spill buffers write in chunks, and the last word of a bitmap is read in one piece. The on-disk attribute and section names are constants inParams; the writers’Typesenum is replaced byDictionarySection. SPECIFICATIONS §7.9 now describes the empty store’s actual layout (index groups with only their seeded directories). - Readers: a remote (HTTP) dictionary section no longer builds the sampled tier index on its first search - it downloaded the whole section - and the tier is built outside the search cache’s lock; the decoded front-coded block cache is weight-bounded (large literals keep fewer blocks); the Zstd decompressor is one per thread per JVM instead of per reader and thread;
-verifyprints each store’s format version andBGReaderexposes it. - Core API: named-graph views share the dataset’s reorder statistics (a
GRAPH ?gquery loaded them once per graph);Graph.stream()works (it threw); property functions registered after the first store opened are visible to BeakGraph datasets; promotable read transactions (begin(),Txn.execute) run as reads instead of throwing. The HTTP channel reads ahead on sequential scans (beakgraph.http.readahead), pins the server’s validator and sendsIf-Rangeso an in-place replacement fails loudly, and keeps query strings (presigned-URL signatures) out of messages, logs and the graph’s identity. - Build: no unused dependencies or resolution repositories, dependency pins synced with Jena 6.2.0, value-based
-Dhdf5.ffm=truebackend selection, a rolling log file, regenerated native-image metadata, CI packages the jars and the benchmarks module and can build the native binary.
0.18.0 - format v5 (2026-07)
- Format v5: RDF 1.2 triple terms (
<<( s p o )>>, nested) stored term-exactly in atripleTermscomponent store in the literals section (SPECIFICATIONS.md §7.6a). v3/v4 files read unchanged; v5 files are rejected by 0.17.0 and earlier with an “Upgrade BeakGraph” error. - Format v4:
rdf:dirLangStringbase directions in an optionallangDirsdataset (§7.5.3). Versions up to 0.17.0 silently stored"x"@en--ltras"x"@en; rebuild affected sources - the file cannot reveal the loss. - Composite (cdt:) literals order lexically in the dictionary (term identity is lexical identity; the value comparison was not a total order). Stores built by 0.17.0 or earlier that contain them must be rebuilt.
- Blank nodes inside composite literals are rejected at ingest.
- Six writer engines (
-method 0..5) targeting one format; the disk engines (1/4/5) bound RAM by spilling and need the native HDF5 library. -export,-verify [-deep],-merge, per-file-threads, VoID statistics opt-in (-void/-voidsketch), reader tuning properties.
0.17.0 and earlier
- Format v3: rank/select directory corrected (v1-2 directories are ignored by readers, which fall back to linear bitmap scans).
- HDF5 container replacing the original Apache Arrow / RO-Crate design.
Format v5 design notes
Code comments cite these by name; they record decisions made while adding RDF 1.2 support, so the rationale stays in the repository.
- One classifier: whether a pattern term is concrete, a variable, or a triple term with embedded variables is decided in exactly one place (
BGIteratorMaster); a second, subtly different classification in an iterator once returned every row for a bound variable. - syntaxARQ at the endpoint: the endpoint parses queries with Jena’s ARQ grammar, which already covers SPARQL 1.2 triple terms and the CDT
FOLD/UNFOLDsurface;syntaxSPARQL_12dropsUNFOLD. The conformance cost is re-checked on every Jena upgrade (src/test/resources/w3c/sparql12/README.md). - Interior terms are dictionary members: an IRI, blank node or literal that occurs only inside a triple term still gets a dictionary id, so the component store can reference it; the writers collect interior terms during the same pass as top-level terms.
- Component resolution: triple terms are stored as component ids into the entities, predicates and literals sections and resolved recursively; the disk engines resolve components with a reference join over the sorted term runs rather than an in-memory map.
- No widened structs: the fixed-stride component record is never widened for a special case (repeated variables, nesting); such cases are handled by the matcher, keeping every engine’s writer identical.
- Mirror topology hazard: six engines implementing one rule by hand is the recurring source of divergence; shared helpers (
RdfSources.parser,Params.gridGraph,RelativeIris,WriterEnginesin tests) exist to keep the rule in one place.