This repository contains the deployment infrastructure for the croissant-live semantic database, leveraging QLever.
This development is funded by:
- The Climate-Adapt4EOSC project has received funding from the Horizon Europe Framework Programme under grant agreement N° 101188248.
- CDIF4EOSC: Developing and implementing the Cross-Domain Interoperability Framework for EOSC is funded by the European Union under Grant Agreement 101292473.
- qlever-tests: The original
qlever-testsrepository is linked here as a Git Submodule. This provides access to the required environment variables (.envfiles) and Docker build contexts (Dockerfiles). - compose.yaml: The Docker Compose file defines the
croissant-liveprofile to run the server and UI concurrently with standard QLever components.
-
Ensure the submodule is fully initialized:
git submodule update --init --recursive
-
Start the infrastructure:
docker compose --profile croissant-live up -d
By default, the Docker compose configuration uses volumes located at ./qlever-tests/volumes and raw data from ./qlever-tests/data.
You can override these directories by providing environment variables before running docker-compose:
VOLUME_DIR=/path/to/my/volumes DATA_DIR=/path/to/my/data docker compose --profile croissant-live up -dThe infrastructure supports running multiple independent instances of Semantic Croissant concurrently on the same machine by using isolated profiles. Environment variables (such as port mappings and DATA_DIR) for each profile are stored in the profiles/ directory (e.g., profiles/codata.env, profiles/cdif4eosc.env, profiles/ca4eosc.env).
To build and run a specific profile (for example, ca4eosc), you must pass the --env-file flag pointing to the profile's configuration file and set the PROFILE_NAME variable.
-
Build the profile-specific API image:
docker build -t api-ca4eosc api/
-
Start the profile using Docker Compose:
PROFILE_NAME=ca4eosc docker compose --env-file profiles/ca4eosc.env -p ca4eosc --profile ca4eosc up -d
Each profile is fully sandboxed, ensuring that QLever index data, Elasticsearch volumes, and Minio objects are kept in dedicated subdirectories based on the profile name (e.g., ./qlever-tests/volumes/ca4eosc/ and ./qlever-tests/data/ca4eosc/).
Before the croissant-live index can be built, the raw Croissant JSON-LD files must be converted into a continuous NTriples format (.nt) suitable for QLever ingestion.
The conversion pipeline script convert_all.py (located in the pipeline/ directory) uses multiprocessing to quickly transform the source JSON-LD files.
If you have a directory of raw Croissant JSON-LD files (e.g., in ../croissant), you can run the pipeline to generate a consolidated data.nt file ready for ingestion:
python3 pipeline/convert_all.py ../croissant ./qlever-tests/data/data.ntThis will produce data/data.nt, which is automatically mounted and read by the server-croissant-live container when indexing begins.
If you run the pipeline again to produce a new or updated data.nt file, you must force QLever to rebuild the internal graph index. QLever will not automatically re-index if the old index cache files still exist in the server volume.
To update the data and restart the services:
- Ensure your updated
data.ntfile is placed in the configured data directory (by default,./qlever-tests/data/data.nt). - Bring down the currently running services:
docker compose --profile croissant-live down
- Remove the old QLever index cache files. Because these files are created by the Docker root user, you can use a transient Alpine container to delete them cleanly from the volume folder:
docker run --rm -v $(pwd)/qlever-tests/volumes/croissant-live/server:/server alpine sh -c "rm -f /server/croissant.*"
- Restart the services:
docker compose --profile croissant-live up -d
Upon startup, the server will detect that the index files are missing and automatically parse the new data/data.nt file to rebuild the index from scratch.
If you already possess the pre-built QLever index files (such as croissant.index.*, croissant.vocabulary.*, etc.) from another machine or backup, you can entirely skip the lengthy index building phase.
- Ensure the
server-croissant-livecontainer is stopped. - Copy all your pre-built index files directly into the mounted server volume directory. By default, this is:
(If you customized
./qlever-tests/volumes/croissant-live/server/
$VOLUME_DIR, place them in$VOLUME_DIR/croissant-live/server/) - Ensure the files are named correctly (e.g., prefixed with
croissant.) and belong to the correct permissions (the QLever Docker container reads them as user65534:0). - Run
docker compose --profile croissant-live up -d. The server will instantly detect the existing index and load the graph into memory.
The Semantic Croissant stack includes a FastAPI service that exposes endpoints for dynamically ingesting new Croissant JSON-LD data into the QLever triple store. The API runs by default on port 7013.
You can add a new Croissant JSON-LD dataset using the /add_record POST endpoint. By default, this will append the converted data to the persistent data.nt file and attempt a live INSERT DATA query to the running QLever instance, making the data instantly queryable without downtime.
curl -X POST "http://localhost:7013/add_record" \
-H "Content-Type: application/json" \
-d @my_dataset.jsonIf you want to additionally trigger a full offline index rebuild (which is useful if the live insertion fails or you want to ensure total consistency), you can pass the rebuild=true query parameter:
curl -X POST "http://localhost:7013/add_record?rebuild=true" \
-H "Content-Type: application/json" \
-d @my_dataset.jsonYou can also trigger a manual index rebuild directly without adding new data using the /rebuild endpoint:
curl -X POST "http://localhost:7013/rebuild"The repository includes the Croissant Toolkit as a git submodule. This Gemini-powered toolkit can automatically generate, enrich, and translate Croissant metadata from raw data or web pages.
You can use the toolkit to generate a dataset and instantly ingest it into QLever:
- Generate Metadata: Use the toolkit's Wizard or Croissant Expert to generate a
.jsonldfile.export GEMINI_API_KEY="your-api-key" python3 croissant-toolkit/.gemini/skills/wizard/scripts/wizard.py "https://example.com/dataset" "My Dataset"
- Ingest into QLever: Once the toolkit generates the
dataset.jsonld, use the Semantic Croissant API to ingest it:curl -X POST "http://localhost:7013/add_record" \ -H "Content-Type: application/json" \ -d @croissant-toolkit/data/croissant/dataset.jsonld
You can directly ingest content from a URL (such as a YouTube video transcript or a web article), convert it into Croissant format, and test the semantic accuracy using our built-in scripts.
Use the url_to_croissant.py converter to download content, generate metadata, slice it into manageable chunks, and perform a dry-run ingestion into the Ollama model.
# Example: Ingesting a YouTube Video
OLLAMA_HOST="http://10.147.18.37:11434" python3 convertors/url_to_croissant.py "https://www.youtube.com/watch?v=OcufTCr3RQs" --sliceThis script will produce a _croissant.jsonld file containing all the sliced semantic chunks, along with their text content and generated summaries.
You can also pass a Google Sheets URL containing a list of URLs to ingest them in bulk. The script will fetch the spreadsheet, extract all valid links, and process them in parallel using the --workers flag.
# Example: Batch ingesting URLs from a Google Sheet using 10 concurrent workers
OLLAMA_HOST="http://10.147.18.37:11434" python3 convertors/url_to_croissant.py --workers 10 "https://docs.google.com/spreadsheets/d/1g8rqbLssGL7lDOGj5nwgXlyDqxrWOQim/edit?gid=573917078#gid=573917078"Once the dataset is converted, you can run the QA accuracy script to automatically ask questions about the text segments and evaluate the answers (with precise provenance).
# Example: Running the QA evaluator
OLLAMA_HOST="http://10.147.18.37:11434" python scripts/test_qa_accuracy.py my_dataset__croissant.jsonldThe script will loop through the ingested chunks, generate contextual questions, answer them, and evaluate the response's Accuracy and Precision on a 1-5 scale.
The repository includes a dedicated MCP service (mcp-croissant-live) that exposes the Semantic Croissant index to AI assistants like Claude Desktop or Cursor.
This enables you to ask AI assistants to "search for datasets about X" or "extract the full JSON-LD for dataset Y", and they will seamlessly execute these tasks against your live QLever instance.
The MCP service is automatically included when you start the croissant-live profile. By default, it runs as an HTTP Server-Sent Events (SSE) server exposed on port 7070.
docker compose --profile croissant-live up -dYou can connect remote MCP clients directly to http://localhost:7070/sse.
If you make changes to the MCP server code (api/mcp_server.py), you need to rebuild the API Docker image and restart the MCP container:
docker build -t api-croissant-live api/
docker compose --profile croissant-live up -d --force-recreate mcp-croissant-liveThe MCP Server integrates with the CODATA ODRL infrastructure for decentralized identity management. This ensures that any Croissant datasets or summaries you save to the Vault are properly attributed to your Decentralized Identifier (DID).
To enable ODRL Authentication:
- Ensure the server is running (
docker compose --profile croissant-live up -d). - Navigate to
https://mcp.dev.codata.org/in your browser. This root page acts as your Authentication Dashboard. - Click either the Google or GitHub OAuth buttons to securely redirect to the external ODRL Wallet (
https://odrl.dev.codata.org/vcs). - Once authenticated, an authorization token (
~/.odrl/authorize) containing your DID is saved to your local machine. - The MCP containers automatically mount this
~/.odrldirectory. Any subsequent AI agent commands that export Croissant JSON-LD or save to the Vault will detect it and uniquely set your DID in thecreatorfield.
The MCP Server implements cryptographic and structural verification of documents saved to the Vault to ensure transparent provenance tracking.
- Digital Signatures & DID Anchors: When a dataset or summary is saved to the Vault via the
save_to_vaulttool, the system automatically computes the UNF-6 hash of the document and generates a Digital Signature combining your authenticated DID and the hash. This signature and a dedicatedserviceverification block are embedded directly inside the.jsonldmetadata, tying the identity of the creator cryptographically to the exact content of the file. - Provenance Verification Tool: You can verify the integrity and provenance of any document in the Vault using the
verify_document_provenanceMCP tool. By supplying the Vault filename (e.g.,kkFU1poLzxYuGowgjjIxYw.md), the tool will read the associated.jsonldmetadata to:- Validate the presence of the DID Verification Block and Digital Signature.
- Extract and list all Creators (Human Users and AI Models) involved in the generation of the document, ensuring complete transparency of the AI/Human collaboration process.
You can connect your IDEs to the public endpoint at https://mcp.dev.codata.org/mcp using the configurations below.
Cursor
Create or edit your MCP configuration file at ~/.cursor/mcp.json or configure it directly through the Cursor Settings UI:
{
"mcpServers": {
"croissant-mcp": {
"type": "sse",
"url": "https://mcp.dev.codata.org/mcp"
}
}
}Windsurf
Edit your configuration file at ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"croissant-mcp": {
"serverUrl": "https://mcp.dev.codata.org/mcp"
}
}
}Zed
Because Zed natively expects stdio for context servers, it requires the mcp-remote proxy bridge to connect to remote SSE servers. Add this to your Zed settings.json:
{
"context_servers": {
"croissant-mcp": {
"command": {
"path": "npx",
"args": ["-y", "mcp-remote", "https://mcp.dev.codata.org/mcp"],
"env": null
}
}
}
}Claude Desktop (Remote)
Claude Desktop also defaults to stdio connections. You can use the mcp-remote proxy bridge in your claude_desktop_config.json to connect to the remote server:
{
"mcpServers": {
"croissant-mcp": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://mcp.dev.codata.org/mcp"]
}
}
}If you are running the semantic-croissant stack locally via Docker, clients like Claude Desktop or Zed can bypass the network entirely and connect directly to the running container via standard input/output (stdio).
To do this, use the following stdio execution command in your client's configuration:
{
"mcpServers": {
"croissant-local": {
"command": "docker",
"args": [
"exec",
"-i",
"semantic-croissant-mcp-croissant-live-1",
"python",
"/app/mcp_server.py"
]
}
}
}Restart your IDE or Claude Desktop. You should see an icon or status indicating the tools are loaded, enabling conversational queries directly against the Croissant datasets!
The Semantic Croissant ecosystem includes a dedicated object storage layer (powered by MinIO) referred to as the Vault. The Vault is designed to persist user interactions, AI-generated dataset summaries, and their corresponding Croissant JSON-LD metadata.
When you use the save_to_vault tool through the MCP server, the content is securely saved to the Vault using a deterministic fingerprint known as the UNF-6 (Universal Numeric Fingerprint) label.
To guarantee data integrity and version control, every saved file is uniquely identified by its contents. Instead of relying purely on random UUIDs, the system generates a UNF-6 hash:
- All words within the text content are split and lexicographically sorted to neutralize minor formatting changes.
- The sorted tokens are concatenated and hashed using SHA-256.
- The resulting hash is truncated to 128 bits and encoded in Base64 (with URL-safe replacements).
The resulting filename adheres to the following structure:
[prefix]_UNF-6_[hash]_[username]_[timestamp].[ext]
For example, saving a dataset summary might produce:
shrink_swell_risks_UNF-6_Kq4bhbZB2z4Vz5lgSYzALA_anonymous_20260810_083505.md
This guarantees that two identical pieces of content will generate the same UNF-6 hash segment, allowing the system to easily track exact duplicates or iterations of a dataset over time. Both the unstructured Markdown (.md) and the structured Croissant metadata (.jsonld) are saved side-by-side using the same naming convention.
The Semantic Croissant stack implements the FAIR Signposting Profile (Level 1) to enhance the machine-actionability of all scholarly objects stored in the Vault.
When retrieving documents from the Vault, the API automatically injects HTTP Link headers containing persistent identifiers, metadata endpoints, and object typing. This enables automated agents and bots to intelligently traverse the scholarly web without needing to scrape HTML or parse ad-hoc formats.
You can verify the presence of the Signposting headers using a simple GET request. The Link header acts as a map for machines:
curl -v http://localhost:7070/vault/honduras_president_charges_factual_summary.md > /dev/nullExpected Output:
< HTTP/1.1 200 OK
< link: <https://mcp.dev.codata.org/vault/...>; rel="cite-as", <https://mcp.dev.codata.org/vault/...jsonld>; rel="describedby" type="application/ld+json", <https://mcp.dev.codata.org/vault/...>; rel="item" type="text/markdown", <https://schema.org/Dataset>; rel="type", <https://creativecommons.org/licenses/by/4.0/>; rel="license"
< x-fair-signposting: enabled
< content-type: text/markdown; charset=utf-8Since the MCP server acts as an intelligent proxy to the vault, AI agents inherently leverage these endpoints when reading articles using the read_vault_article tool. You can test this locally by querying the tool directly via the docker container:
docker exec semantic-croissant-mcp-croissant-live-1 python3 -c "
import asyncio, sys, os
sys.path.append(os.path.join(os.getcwd(), 'api'))
from mcp_server import call_tool
asyncio.run(call_tool('read_vault_article', {'url_or_filename': 'honduras_president_charges_factual_summary.md'}))
"The MCP Server includes an optional feature to automatically upload generated Croissant (.jsonld) and Markdown (.md) files directly to a Google Drive folder. The folder is automatically named after the user's email ID.
To enable this feature, you must configure a Google Service Account:
- Go to the Google Cloud Console and create a new project (or use an existing one).
- Enable the Google Drive API for your project.
- Navigate to IAM & Admin > Service Accounts and create a new Service Account.
- Create a new JSON key for this Service Account and download it to your local machine.
- Save the file as
credentials.jsonin the root of this project (it is mapped into theapi-croissant-livecontainer via thecompose.yamlvolumes).- Alternatively, you can specify a custom path using the
GDRIVE_CREDENTIALS_FILEenvironment variable.
- Alternatively, you can specify a custom path using the
When calling the url_to_croissant tool from your AI assistant, simply pass the upload_gdrive: true parameter. The server will authenticate using the mapped credentials.json and upload the extracted files into a Drive folder matching your ODRL/Authentication email address! If the credentials file is missing, the upload is gracefully skipped.