Help
About THEOBROMA
THEOBROMA is an open, curated natural products database aggregating compounds from 29 source databases across four biological kingdoms (plant, fungi, bacteria, animal), a residual unresolved category, and 13 resolved geographic regions. Cross-kingdom attributions are preserved via a secondary-kingdoms layer so a single compound can be discovered under any kingdom in which it has been reported. THEOBROMA supports structure-based virtual screening, cheminformatics research, and drug discovery from natural product chemical space.
Search
Keyboard shortcut: Press / anywhere on the site to focus the search box without clicking.
Compound name — partial substring matching, case-insensitive. Autocomplete suggestions appear after 2 characters. Also searches PubChem synonyms (e.g., "alpha-lipoic acid" finds "Thioctic acid"). Hyphens and Greek letters are normalized automatically. The "Exact match" toggle restricts to literal equality on the chosen field. Use the "/" keyboard shortcut to focus the search box from anywhere on the page.
SMILES — exact canonical SMILES match. The query SMILES is canonicalized via RDKit before lookup, so equivalent representations resolve to the same compound.
InChIKey — exact 27-character match (full stereochemistry-aware key). If the compound is found, the search redirects directly to the compound detail page.
Source organism — substring matching against the producing organism field with autocomplete.
Source database and Kingdom — case-insensitive matching against controlled vocabularies (29 databases; four resolved kingdoms plus an unresolved category: plant, fungi, bacteria, animal, unresolved). Filtering by kingdom returns every compound attested in that kingdom (primary or secondary) so cross-kingdom metabolites surface from each of their kingdoms.
Geographic region — each compound carries a set of macro-regions (13 resolved regions), not a single value. Filtering by a region returns every compound attributed to that region, so a compound that occurs across several regions surfaces under each of them (for example, a pantropical compound is found under both Africa and Southeast Asia). Region membership is derived by combining the native geographic distribution of a compound's source organisms (World Checklist of Vascular Plants, WCVP) with the curated scope of region-specific source databases; see "Geographic regions" below.
Genus, Family, Order, Class, and Phylum — taxonomic search across all five ranks. Resolution combines three authorities: the World Checklist of Vascular Plants (WCVP, Kew, January 2026 release) for plant binomials, NCBI Taxonomy (taxdump 2026-05) for non-plant taxa, and the World Register of Marine Species (WoRMS) for marine organisms. Kingdom is resolved for essentially all compounds (~100%), while ranks below kingdom (phylum, class, order, family, genus) are resolved for roughly 29% of compounds: a compound's full lineage is recovered only when its source organism matches a backbone entry, so these deeper ranks resolve together (all near 28-29%) or not at all. Coverage is highest for organisms with clean binomials in WCVP/NCBI/WoRMS and lowest for compounds whose source records give only vague or aggregate organism strings. Genus and family come primarily from WCVP for plants and NCBI for non-plants; order, class, and phylum come from NCBI and WoRMS JSONB lineages. Synonyms are automatically mapped to accepted names. Note that NCBI Taxonomy does not consistently fill order for plant clades, so order-rank queries for plant taxa may return fewer results than expected; family or genus search is recommended for plant queries.
Chemical property — filter by chemical class or by numerical property ranges. Two class-filter types are available reflecting the two underlying classification ontologies:
- NPClassifier class — queries the
np_classandinferred_classfields together. NPClassifier is the natural-products-focused ontology; assignments are inherited from source databases where present and computed via the canonical GNPS2 NPClassifier service otherwise.inferred_classcovers compounds where the NPClassifier output was empty, filled by a per-class XGBoost ensemble trained on Morgan + MACCS + ChemBERTa features. - ClassyFire superclass — queries the
classyfire_superclassfield, carrying forward ClassyFire's broad chemical-taxonomy assignment where the source database provided it.
Numerical property ranges are also available for exact mass, LogP, TPSA, HBA, HBD, rings, rotatable bonds, plus 41 ADMET-AI endpoints (toxicity, absorption, metabolism, distribution and physicochemical predictions including AMES, hERG, DILI, BBB, bioavailability, P-gp, Caco-2, clearance, half-life, plasma protein binding, solubility, lipophilicity, volume of distribution, the seven Tox21 nuclear-receptor assays, the five Tox21 stress-response assays, and CYP1A2/2C9/2C19/2D6/3A4 substrate and inhibition).
Behavior notes for class search. Class search uses case-insensitive substring matching against the underlying classifier columns (np_class, classyfire_superclass, inferred_class, effective_superclass, effective_pathway). Queries for labels not present in any of these columns return zero rows. For example, "Curcuminoids" is not used as a class by NPClassifier or ClassyFire as applied to this corpus, so a class-filtered query for "Curcuminoids" returns no compounds; to retrieve curcumin and structural analogs, use name search (q=curcumin&type=name), which substring-matches against compound names and PubChem synonyms.
Name-search behavior. Name search returns substring matches against compound names and synonyms. A query such as q=curcumin matches any compound whose name or synonyms contain "curcumin" as a substring, including derivatives, conjugates, and analogs that list curcumin in their synonyms. To restrict to literal-equality name matches, enable the "Exact match" toggle.
Multi-filter search
Click "+ Add another filter" to combine multiple criteria with AND logic. For example, search for all triterpenoids from Africa, or all fungal compounds with exact mass 200-500 Da and low DILI risk.
Results are relevance-sorted for name searches: exact matches appear first, then prefix matches, then substring matches. Within each tier, shorter (simpler) names rank higher. All result tables have sortable columns — click any column header to sort ascending or descending.
Similarity search
Enter a SMILES string, compound name, or THEO_ ID to find structurally similar compounds. Four metrics are available:
- Morgan / ECFP4 Tanimoto (default) — circular fingerprint, radius 2, 2048-bit. Good for general structural similarity.
- MACCS Tanimoto — 167-bit substructure key fingerprint. Good for matching based on functional-group composition; less sensitive to large skeletal differences.
- ChemBERTa cosine — 768-dimensional dense embedding from a transformer pretrained on ZINC. Captures learned chemical-space features that may differ from hand-engineered fingerprints.
- NaFM cosine — 1024-dimensional embedding from a graph neural network pretrained specifically on natural products (contrastive learning plus masked-graph modeling). Unlike ChemBERTa (trained on synthetic molecules from ZINC), NaFM captures scaffold-derived and biosynthetic patterns native to natural-product chemical space, so nearest neighbours tend to be biosynthetically related analogs. Available for compounds resolvable to the corpus (by THEO_ id, exact name, or InChIKey); novel off-corpus SMILES are not embedded at query time.
When to use which: Morgan/ECFP4 is the default for general-purpose similarity ranking and is what most cheminformatics tools use. MACCS is preferred when functional-group composition matters more than overall skeleton (e.g. looking for compounds with similar pharmacophore patterns despite different scaffolds). ChemBERTa is useful when you want learned representations that may pick up subtle features fingerprints miss. NaFM is preferred when you want neighbours related in a natural-product-aware sense (shared scaffold or biosynthetic origin), since it was pretrained on natural-product chemical space rather than synthetic libraries.
Threshold (Tc) defaults to 0.3; Morgan and MACCS apply this as a hard cutoff. The embedding metrics (ChemBERTa and NaFM) instead use FAISS indexes and return the top-N nearest neighbours: ChemBERTa via an HNSW index (M=32, efConstruction=200, efSearch=128), NaFM via a memory-efficient fp16 flat index over the corpus. Results can be post-filtered by region, kingdom, and chemical class.
Substructure search
Enter a SMILES or SMARTS pattern to find all compounds containing that substructure. Uses Morgan fingerprint pre-screening followed by RDKit substructure matching.
Scaffold browser
Browse compounds grouped by their Bemis-Murcko molecular framework. Click any scaffold to see all compounds sharing that core structure.
Property computation provenance
Physicochemical properties (exact mass, LogP, TPSA, HBA, HBD, ring count, rotatable bonds) are computed by THEOBROMA from canonical SMILES via RDKit (Descriptors.ExactMolWt, Descriptors.MolLogP, Descriptors.TPSA, rdMolDescriptors functions). Source-database property values are not carried forward, ensuring consistent computation across the corpus.
Chemical classifications come from two ontologies. NPClassifier (NP-class, NP-superclass, NP-pathway) is inherited from source databases where present and computed via the canonical GNPS2 NPClassifier service otherwise. ClassyFire superclasses are carried forward where present in source databases; for compounds lacking a ClassyFire annotation, the inferred_class field provides a per-class XGBoost prediction trained on Morgan, MACCS, and ChemBERTa features. No acceptance threshold is applied at assignment; the per-compound top-class probability is persisted so a threshold can be imposed downstream. NPClassifier is trained on organic natural-product chemical space and does not reliably classify organometallic or inorganic structures; 170 single-fragment transition-metal compounds retain a classification that should be treated as low-confidence.
ADMET predictions (41 ADMET-AI endpoints persisted alongside six RDKit-computed columns, covering toxicity, absorption, metabolism, distribution, and physicochemical properties) are produced by ADMET-AI (Chemprop D-MPNN ensemble; Swanson et al.). Each endpoint has its own min/max/mean range visible in the search dropdown and is queryable via URL parameters of the form <endpoint>_min and <endpoint>_max. Note on applicability: ADMET-AI's models are trained on drug-like molecules (Therapeutics Data Commons), so predictions for large or structurally complex natural products lie outside the models' training domain and should be read as indicative rather than quantitative — roughly 26,000 compounds in the corpus exceed 1000 Da, well beyond typical drug-like chemical space.
For the 19,508 compounds with a matching ChEBI 3-stars entry, additional curator-written metadata is surfaced on the compound detail page: a chemistry definition, an IUPAC name (where ChEBI's IUPAC variant differs from the THEOBROMA primary name), PubMed PMID arrays, and per-database deep-links to the external resources cross-referenced by ChEBI. Matching is by InChIKey; no transformation of the THEOBROMA structure or naming is performed.
Classification
Every compound receives a chemical class through a three-tier provenance hierarchy that records where the assignment came from rather than how it was produced. The tier assigned to each compound is shown on its detail page (for example, "Curated").
- Tier 1 — Curated (626,601 compounds) — the class is inherited directly from the source database.
- Tier 2 — NPClassifier (206,042) — the class is assigned by the canonical GNPS2 NPClassifier service where no curated label exists.
- Tier 3 — Model-inferred (300,141) — the class is predicted by a per-class XGBoost model (cross-validated macro-F1 0.81) for compounds without a curated or tool-assigned label; the
inferred_classfield records this together with its source tag. - 21 compounds could not be featurized and remain unclassified.
Superclass and pathway are derived from the class through the NPClassifier ontology, so they remain consistent across all three tiers. Pathways use the seven canonical NPClassifier categories, and a compound may belong to more than one. This ordering is a provenance hierarchy (which source assigned the label), not a confidence ranking.
Multi-source compounds
When a compound (identified by full 27-character InChIKey) is reported by more than one source database, THEOBROMA stores it as a single row with all contributing databases listed in the all_sources field (pipe-separated). The primary source_db preferentially shows regional or specialized databases over aggregators like COCONUT. Compounds appearing in two or more source databases are counted in the "Multi-source compounds" statistic on the Statistics page.
Compounds shown as "global / unresolved" for region carry no resolved macro-region — their source organisms have no WCVP native-distribution record and they come from no region-specific database; likewise "unresolved" kingdom means no taxonomic metadata was available. Geographic distribution is available for 264,039 compounds; the remaining corpus is global/unresolved.
Compounds reported across multiple kingdoms (e.g. a plant database and a fungal database for the same metabolite, or an endophyte metabolite isolated from both host plant and fungus) are stored as a single row with a primary kingdom (the majority-attested one) and a secondary_kingdoms array listing every additional kingdom in which the compound is reported. Filtering by kingdom returns every compound attested in that kingdom regardless of whether it is primary or secondary, so endophyte metabolites surface from the plant kingdom as well as the fungal kingdom. The corpus total (1,132,805 distinct compounds) remains consistent across views; sums of kingdom slices exceed the corpus total only because cross-attributed compounds contribute to several slices.
Geographic regions
Geographic region is multi-valued: a compound belongs to every macro-region in which it is attributed, not one. The region set for each compound is the union of two signals — (i) the native distribution of its source organisms, taken from WCVP and mapped from finer botanical regions onto 13 macro-regions, and (ii) the curated geographic scope of region-specific source databases (for example ANPDB, AfroDb and SANCDB for Africa; IMPPAT for South Asia; TIPdb-3D for East Asia; CSIRO for Australia; NaturAr for Latin America). Global aggregator databases contribute geography only through their organisms' WCVP distribution, never a fixed tag. Of 1,132,805 compounds, 264,039 carry at least one region and 136,047 span more than one; the earlier single-region field, which forced one arbitrary value per compound, has been replaced by this set. Because compounds can belong to several regions, region counts and the world map sum beyond the region-covered total — the same multi-attribution behaviour as kingdoms. On the world map, a country shared by two macro-regions (for example the Central Asian states, which are both Central Asia and Russia/CIS) is shaded by its stronger region rather than the sum.
How to cite
Thor Klamt, Anne Jaczkowski, Jakob Franke and Wolfgang Nejdl (2026). THEOBROMA: an aggregated open database of 1.13 million natural products with per-compound license auditing, three-tier classification, and stereochemistry-aware deduplication. Preprint: https://doi.org/10.64898/2026.06.12.731585 Dataset: https://doi.org/10.5281/zenodo.22816330
Please also cite the original source databases for any compounds you use. Source database citations and licenses are listed on the Statistics page.
API
GET /api/search?q=QUERY&type=TYPE&limit=N&offset=M&format=FORMAT&license=LIC
Parameters:
q — search query (required)
type — name | smiles | inchikey | kingdom | organism |
region | source | genus | family | order | tax_class | phylum |
npclassifier_class | npclassifier_superclass | classyfire_class |
class (union of chemistry-class ontologies) | pathway |
property (concrete field given in prop_type)
prop_type — with type=property, the concrete field: npclassifier_class,
npclassifier_superclass, classyfire_class, pathway, or a
taxonomic rank
exact — true to require literal equality on text-search types
(default false; substring match)
limit — max results, 1-10000 (default: 50)
offset — pagination offset (default: 0)
format — json | csv (default: json)
license — all | commercial | academic (default: all)
GET /api/compound/COMP_ID — full compound details by THEO_ id (includes license_provenance sub-object)
GET /api/compound/COMP_ID/license-provenance — per-source license attestation chain (audit trail)
GET /api/stereoisomers/COMP_ID — all stereoisomers (14-char InChIKey match)
GET /api/similarity?smiles=...&metric=morgan|maccs|chemberta|nafm
GET /api/autocomplete?q=QUERY — name autocomplete
GET /api/organisms?q=QUERY — organism autocomplete
GET /api/filter_options — available values per filter type
(includes npclassifier_class, classyfire_class)
GET /api/search?...&format=json&download=true — download results as JSON attachment
GET /api/bulk?cols=...&format=json — bulk download in JSON format (streamed)
GET /api/stats — database statistics (includes license-tier breakdown)
GET /api/bulk?cols=...&license=commercial|academic|all — streaming bulk export (CSV, default 14-column annotation payload); tier=open|nc|all also accepted
Full OpenAPI specification at /api.
Recommended bulk-download size: up to 100,000 rows per request through /export or /api/search; use /api/bulk for full-corpus exports (streaming, no row limit).
No per-request rate limit is enforced at present; for sustained high-throughput use (more than a few requests per second), please contact the corresponding author so we can plan server load. Known scraping user-agents (GPTBot, CCBot, ClaudeBot, anthropic-ai, PerplexityBot, Bytespider) are rate-limited on the heaviest routes; robots.txt documents the policy.
License
Web application code: MIT License.
Compound data: per-record license_tier describing the terms under which a compound's chemical structure may be redistributed, derived from source databases by a two-level rule.
Resolution rule. First, each source database is assigned the most restrictive license applicable to its contents: where a source is non-commercial, share-alike, or its terms cannot be resolved at the per-compound level, the whole source is treated conservatively at that tier (or as Unspecified when no license can be determined). Second, a compound found in multiple sources takes the most restrictive license among them, so that redistribution terms hold for every source that reports the structure. The assignment is therefore conservative throughout: a compound is only placed in a permissive tier when every source attesting it permits those terms. Tier order (permissive to restrictive): CC0 < CC BY 4.0 < CC BY-NC 4.0 < CC BY-NC-SA 4.0 < CC BY-NC-ND 4.0 < Unspecified.
Each source is assigned a single license tier; a compound's tier is then the most restrictive across all sources that report it, so a source's compounds may appear in a more restrictive tier than its own assignment when co-attested elsewhere.
Per-source assignment.
- CC BY 4.0: COCONUT 2.0, LOTUS v11, MIBiG 3.1, SANCDB, CMAUPv2, AMDB, NaturAr, AfroDb, CyanoMetDB, EMNPD, MeFSAT.
- CC BY-NC 4.0: NPASS 3.0, FooDB, NPAtlas, HERB 2.0, IMPPAT 2.0, TIPdb-3D, ANPDB, CSIRO, phytochemdb, Phyto4Health, StreptomeDB, MycoCentral.
- CC0: TM-MC 2.0.
- Unspecified: MicotoXilico, LMDB (Lichen), and specialized collections (CMDB Cereals, SMDB Spice, TMDB Trichoderma) whose license could not be resolved at the per-compound level.
Resulting per-compound distribution: CC BY 4.0 — 887,999 (78.39%); CC BY-NC 4.0 — 229,397 (20.25%); CC0 — 8,243 (0.73%); Unspecified — 7,166 (0.63%). Unspecified rows are excluded from both "Commercial use allowed" and "Non-commercial research" filters pending source-license confirmation.
896,242 compounds (79.12%) permit commercial use (CC0 or CC BY 4.0). The license_tier field describes the structure's redistribution terms; redistributing the full integrated annotation record for a compound (source organism, literature references, computed metadata) may additionally require honoring the contributing sources' terms for those annotation fields. The CMNPD database, originally part of the corpus, was removed in v32 due to its CC BY-NC-SA share-alike clause; CMNPD-exclusive compounds were dropped and CMNPD provenance was stripped from multi-source rows.
Glossary
Compact reference for non-cheminformatician users.
- SMILES
- Simplified Molecular-Input Line-Entry System; an ASCII string representation of a molecular graph (e.g.
CC(=O)Oc1ccccc1C(=O)Ofor aspirin). - SMARTS
- A SMILES-derived pattern language for substructure queries; supports wildcards and logical operators.
- InChI / InChIKey
- InChI is a normalized textual identifier capturing the full molecular graph and stereochemistry. The 27-character InChIKey is its compact hash, suitable for exact-match lookup.
- Canonical SMILES
- A SMILES string produced by a deterministic canonicalization algorithm so that the same molecule always yields the same string regardless of input order.
- Tanimoto coefficient (Tc)
- Similarity metric between two binary fingerprints, computed as the size of their intersection divided by the size of their union. Range 0 to 1; higher means more similar.
- Morgan / ECFP4 fingerprint
- Circular fingerprint encoding atom-environment substructures within radius 2 (ECFP4 convention). Folded to 2048 bits in THEOBROMA. The cheminformatics community default for similarity ranking.
- MACCS keys
- 167-bit substructure-presence fingerprint with a curated dictionary of functional-group and ring-system queries. Good for functional-group similarity; weaker for whole-skeleton similarity.
- ChEBI
- Chemical Entities of Biological Interest; a manually-curated chemistry ontology maintained by EMBL-EBI under CC BY 4.0. THEOBROMA integrates the 3-stars (highest curation tier) subset, matching by InChIKey. For 19,508 compounds, ChEBI provides a curator-written definition, IUPAC name, PubMed PMID arrays, and cross-references to KEGG, KNApSAcK, LIPID MAPS, MetaboLights, ChEMBL, MetaCyc, UniProt, Rhea, NMRShiftDB, CompTox, and other resources.
- ChemBERTa embedding
- 768-dimensional dense vector produced by a transformer model pretrained on a 100k-molecule subset of ZINC. Captures learned chemical-space structure; queried via cosine similarity in a FAISS HNSW index for sub-second retrieval.
- NaFM
- Natural-products Foundation Model; a graph neural network pretrained on natural-product structures (contrastive and masked-graph objectives) that produces a 1024-dimensional embedding capturing NP-native scaffold and biosynthetic features. Used in THEOBROMA as a fourth similarity metric, complementary to the synthetic-molecule-trained ChemBERTa; retrieved via a memory-efficient fp16 FAISS index over the corpus.
- Bemis-Murcko scaffold
- The molecular framework obtained by stripping all acyclic side chains while retaining linkers between ring systems. Used to group compounds by core skeleton.
- FAISS HNSW
- Hierarchical Navigable Small World graph index for approximate-nearest-neighbor search over dense embeddings. Enables fast top-N retrieval in 1M+ compound corpora.
- NPClassifier
- Deep-learning model that assigns natural-product class, superclass, and biosynthetic pathway from molecular structure (Kim et al., 2021). In THEOBROMA, the
npclassifier_classfilter type queries both the direct NPClassifier output (np_class) and the per-class XGBoost-inferred extensions (inferred_class). - ClassyFire
- Rule-based chemical taxonomy assigning kingdom, superclass, class, and subclass via expert-curated structural patterns (Djoumbou-Feunang et al., 2016). In THEOBROMA, the
classyfire_classfilter type queries the superclass field carried forward from source databases. - ADMET
- Absorption, Distribution, Metabolism, Excretion, Toxicity; properties that predict in-vivo behavior of drug candidates. THEOBROMA persists all 41 endpoints of the ADMET-AI Chemprop ensemble (Swanson et al.) in the searchable database, including AMES, hERG, BBB, bioavailability, P-gp, Caco-2, multiple CYP inhibition and substrate predictions, the Tox21 nuclear-receptor and stress-response panels, clearance, half-life, plasma protein binding, solubility, lipophilicity, LD50, and volume of distribution.
- DILI
- Drug-Induced Liver Injury; predicted probability that a compound causes hepatotoxicity.
- AqSolDB solubility
- Predicted aqueous solubility (log mol/L) from a model trained on the AqSolDB benchmark.
- LogP
- Logarithm of the octanol-water partition coefficient; measures lipophilicity. Drug-like compounds typically have LogP between 0 and 5.
- TPSA
- Topological polar surface area in square angstroms; correlates with membrane permeability and oral bioavailability.
- HBA / HBD
- Hydrogen-bond acceptor and donor counts; together with MW and LogP, part of Lipinski's Rule of Five for oral drug likeness.
- Trust score
- THEOBROMA-internal composite (0-1) over nine annotation-coverage axes including structure, classification, scaffold, taxonomy, geography, traditional-medicine tagging, ADMET coverage, novelty, and similarity-to-bioactive scoring. Higher means more confidently characterized.
Source databases (29)
COCONUT 2.0, LOTUS v11, FooDB, NPASS 3.0, HERB 2.0, TM-MC 2.0, IMPPAT 2.0, CSIRO Australian NP, ANPDB, NPAtlas, phytochemdb, MicotoXilico, StreptomeDB, MIBiG 3.1, EMNPD, MeFSAT, CyanoMetDB, MycoCentral, NaturAr, LMDB_Lichen, AMDB, Phyto4Health, AfroDb, SANCDB, CMDB_Cereals, TMDB_Trichoderma, SMDB_Spice, CMAUPv2, TIPdb-3D.
Enrichment layers
In addition to the 29 source databases, THEOBROMA integrates several enrichment layers that annotate compounds with metadata not directly available from any single source:
- LOTUS (via Wikidata SPARQL): natural-product to organism linkages, references, and chemical classifications.
- PubChem synonyms (PUG REST + RDF bulk): compound name variants, CAS, INN, USAN, brand and trade names.
- WCVP (World Checklist of Vascular Plants, Kew): plant genus and family taxonomy. Combined with NCBI Taxonomy and WoRMS, kingdom is resolved for ~100% of compounds and the ranks below kingdom (phylum through genus) for ~29%, after NCBI- and WCVP-based backfilling for plant clades.
- GBIF (Global Biodiversity Information Facility): species-name resolution. (Geographic region is now derived primarily from WCVP native distribution combined with region-specific source databases; see Geographic regions above.)
- ChEBI (Chemical Entities of Biological Interest, EBI; 3-stars manually-curated subset): curator-written definitions, IUPAC names, PubMed PMIDs, and cross-references to KEGG, KNApSAcK, LIPID MAPS, MetaboLights, ChEMBL, MetaCyc, UniProt, Rhea, NMRShiftDB, CompTox, CAS, and BRENDA. ChEBI annotations apply to 19,508 compounds (1.72% of corpus); 84,019 net new synonyms were merged into the searchable synonyms layer. License: CC BY 4.0.
Database update cadence
THEOBROMA tracks source-database releases and re-integrates on a roughly biannual cycle. Per-source download dates and version strings are tracked in sources.yaml. Major releases are tagged in the GitHub repository and deposited on Zenodo for long-term archival.
Visitor logging
Requests are logged for service operation and abuse detection, on the legal basis of Art. 6(1)(e) GDPR in conjunction with § 3 NDSG and § 3 NHG. No IP address is stored: each request carries a token derived from the client address under a key that is generated in memory, rotated daily and never written to disk, so tokens cannot be linked back to an address or across days. A country code resolved from a local table, the request path, the timestamp and the user-agent string are stored alongside it. The site sets no cookies and neither writes to nor reads from your device, so no consent under § 25 TDDDG is required. No analytics or advertising service is used, and the visitor map on the Statistics page is built from country codes only.
Contact
For bug reports and feature requests, please open an issue on GitHub. For other inquiries, contact the corresponding author Thor Klamt at thor.klamt@l3s.de or Thor.Klamt@gmail.com.