Releasing data to Zenodo¶
opsci zenodo publishes a project's data/ to Zenodo as a versioned record. The
open-science-publish:zenodo-release skill calls it after a person confirms the release. Every
published version is permanent and public, so release at publication points, not on every
data change.
Commands¶
opsci zenodo release [ROOT] --dry-run # print the plan; no network, nothing written
opsci zenodo release [ROOT] --version 1.0 # release to the sandbox (the default)
opsci zenodo release [ROOT] --version 1.0 --production # release to zenodo.org
opsci zenodo checksum data/<task-id> [--root ROOT] # sha256 of the reproducible tar of a path
opsci zenodo check-token [--production] # check the token file and that Zenodo accepts it
check-token refuses a missing token file, a token file that gives any permission to group
or others (use mode 600), and a file that holds anything but one token. Then it makes one
read-only request (a list of at most one deposition) to see whether Zenodo accepts the token.
It creates nothing. It takes --api-url and --token-file as release does.
Options of release:
--dry-run: builds each tar in memory to get its checksum and size, then prints the groups, checksums, reuse decisions and limit checks. It makes no network call and writes no file.--production: required for zenodo.org. Without it the tool uses sandbox.zenodo.org, and an--api-urlthat points at zenodo.org is refused. With it the tool publishes a permanent, public record with no confirmation prompt. Thezenodo-releaseskill's instructions are what require the user's explicit confirmation first: an agent must not pass--productionuntil the user has confirmed the version and file list in the conversation.--write-citation: write the DOIs toCITATION.cffand to the datasets indata/MANIFEST.yamlalso for a sandbox release. Production releases always do this. Sandbox DOIs do not resolve, so by default they are recorded only in the manifest'szenodo.sandboxsection.--token-file,--api-url,--build-dir: override the defaults below.
Token¶
Create a personal access token with the scopes deposit:write and deposit:actions:
on sandbox.zenodo.org for testing, on zenodo.org for real releases. These are separate
accounts with separate tokens. Store each in its own file, readable only by you:
mkdir -p ~/.config/opsci
( umask 077; cat > ~/.config/opsci/zenodo-sandbox.token ) # paste, then Ctrl-D
( umask 077; cat > ~/.config/opsci/zenodo.token ) # production
($XDG_CONFIG_HOME replaces ~/.config if it is set.) The tool refuses a token file that
group or others can read. It never takes the token from the command line or from an
environment variable, and never prints it. The token is sent only in the Authorization
header, and only to the host of the API URL: a link in a server response (the new-version
draft, the upload bucket) that points to another host, port or scheme is refused. An
--api-url must use https; plain http is accepted only for a loopback test server
(127.0.0.1, localhost, ::1).
What a release does¶
- Tar groups. Data is packed into tar groups, one
<group>.tar.gzeach. By default each entry ofdata/(normally one directory per task) is one group. To group differently, for example to put finished data apart from data that still changes, list groups indata/MANIFEST.yaml:
zenodo:
groups:
- name: finished
paths: [data/t01-setup, data/t02-first-fit]
- name: t05-sweep
paths: [data/t05-sweep]
Groups may not overlap. Entries of data/ that no group covers are reported and not
released.
Data of private tasks is left out; see Private data. Before any tar is
built, every file to be released is scanned for secrets and leaks; see Scans.
2. File list. FILES.tsv is uploaded next to the tars. It lists every file, one per line,
in six tab-separated columns: path, size, sha256, tar (the tar that holds it),
task (the task id when the file lies under data/<id>/ and <id> is a task of the
project, in tasks/, brainstorm/tasks/ or a verifications/ directory; else empty) and
results (the ids of the result nodes whose artifacts name the file or a directory above
it, comma-separated). The tasks and results come from the node headers (opsci map build
reads the same ones), so FILES.tsv also changes when a header changes. A change to
FILES.tsv alone does not make a new version: with no tar changed, the release is
still refused as "nothing changed".
3. Limits. A Zenodo record holds at most 100 files and 50 GB. The tool counts the groups
plus FILES.tsv before it builds anything, and adds up the sizes after the tars are built.
If either limit is exceeded it stops before any network call.
4. Reproducible tars. Members are sorted by name, with mtime 2000-01-01, owner and group 0,
no user names, and modes normalised to 644/755 (777 for symlinks). The gzip header has no
file name and no timestamp. The same input gives a byte-identical tar and the same checksum,
wherever and whenever it is built. Symlinks are handled as described under
Symlinks. The tars are built with Python's tarfile, so the result does not
depend on the installed tar.
5. Versions. The first release creates a deposition. Each later release asks Zenodo for a
new version of the previous record. Zenodo copies the previous version's files into the new
draft. The tool keeps every file whose checksum is unchanged, deletes groups that no longer
exist, and uploads only groups whose checksum changed, plus FILES.tsv. If nothing changed,
it refuses to make a new version.
6. Checks and publish. Each upload's md5 and size must match the local file, and the draft
must hold exactly the planned files. Then the metadata is set and the draft is published.
If a run stops before publishing, the draft id is kept in the manifest and the next run
continues with that draft (Zenodo allows only one unpublished draft per record).
7. Write-back. The version DOI, the concept DOI (which always resolves to the newest
version), the checksum of every uploaded file and the paths each tar holds are written to
data/MANIFEST.yaml, under zenodo.sandbox or zenodo.production:
zenodo:
production:
releases:
- version: '1.0'
date: '2026-09-28'
record_id: 1234567
doi: 10.5281/zenodo.1234567
files: {t04.tar.gz: {sha256: ..., md5: ..., size: ...}, FILES.tsv: {...}}
groups: {t04.tar.gz: [data/t04]}
concept_doi: 10.5281/zenodo.1234566
For production (or with --write-citation), each dataset inside a released group gets
zenodo: <version DOI>, and CITATION.cff gets two identifiers entries: the concept DOI
and the version DOI.
The tool rewrites only the zenodo: section and the zenodo: field of the datasets, in
place, so comments elsewhere in the manifest (a note after a sha256:, a comment line
between datasets) are kept. Comments inside the zenodo: section are not kept: the tool
writes that section. If a manifest cannot be edited in place (the rewrite would not read
back as the intended data), it is written whole and only its leading comment block is kept.
Symlinks¶
A group root (an entry data/<name>, or a path in zenodo.groups) may be a symlink, for
example data/<task-id> pointing to scratch. The tool follows it and archives what it points
to, if the target is allowed:
- The target must lie inside the project, inside the site's
scratchdirectory, or inside a directory listed underdata_rootsinconfig/site.local.yaml:
scratch: <your scratch directory>
data_roots: [<a directory>] # more directories that data/ symlinks may point into
- Whatever those settings say, the tool refuses a target that is or holds the home directory
or the project, one inside a hidden directory of the home directory (
~/.ssh,~/.config,~/.aws,~/.gnupg, ...), one that is, holds or lies inside the config directory ($XDG_CONFIG_HOMEor~/.config, which holds the token files inopsci/), and one inside the project's.git.
A symlink below a group root is stored as a symlink, not followed. It must be relative and
point to a place inside the same group root (sub/link -> ../a.txt is stored; link ->
~/file and link -> ../../other-task/file are refused). Replace a refused link with a
relative one or with the file itself.
opsci zenodo checksum applies the same rules.
Private data¶
Data of a task whose privacy is soft-private or hard-private is not released. The task's
tier is its privacy field, else policy.default_privacy of publish/manifest.yaml; a
verification task inside a private task is at least as private; a task directory whose
context.md has no valid node header counts as hard-private, as in opsci publish. A path
under data/ that the hard_private list of publish/manifest.yaml names is hard-private
too.
- With the default groups (one per entry of
data/), the entry of a private task is left out and the plan says so. - A group in
zenodo.groupsthat holds a private path is refused.
FILES.tsv and the record description therefore never name a private task's files. To
release a private task's data, the user decides it and names the task in the data manifest:
zenodo:
include_private: [t05-sweep] # private tasks whose data the user has decided to release
Scans¶
Before any tar is built, every file to be released, every name and every symlink target is
scanned with the secret patterns of opsci publish (tools/opsci/secretscan.py) and its leak
patterns (tools/opsci/leakscan.py: absolute paths, emails, IP addresses, SLURM job numbers,
the user, host and config/site.local.yaml values, and the project's
publish/PRIVATE_POLICY.md patterns). The generated FILES.tsv is scanned too. Any finding
refuses the release, in the dry run as well, before any network call. Secrets are shown
redacted.
- A file up to 256 MiB is scanned as
opsci publishscans it: text with every pattern; PNG and JPEG metadata; PDF strings and text; the members of zip (also.npz), tar, gz, bz2 and xz archives (an archive that cannot be read, such as 7z or RAR, refuses the release). One difference: in binary data,absolute-pathruns only on long printable runs (16 or more printable characters), as the email and IP address patterns do. On all of a compressed stream it matches by chance about once per megabyte. A path stored as a string, such as an HDF5 attribute, is such a run and is found. - A bigger file is read in overlapping chunks, with the same patterns; archive members in it are not opened. A release of many gigabytes takes a while to scan; the scan reads every byte once.
gitleaksis not run on data (it is run byopsci publish).
The kinds that opsci publish lets the user accept (SLURM job numbers: slurm-job-id,
slurm-array-id, slurm-out-file) may be accepted here too, in the data manifest, with the
same fields as in publish/manifest.yaml (see the publish skill's check-overrides.md):
zenodo:
overrides:
- check: leak
kind: slurm-job-id
paths: [data/t05-sweep/logs] # optional
reason: Job numbers in run logs name no person or machine.
date: 2026-10-05
The plan lists every accepted finding. A secret, and every other leak kind, is never overridden: change or remove the file, or leave it out of the tar groups.
Results pages¶
After a production release, run opsci map build. Every result artifacts path under
data/ that a production release holds is then shown with that release, in the result's
"stored in" row on the results pages and in the "stored in" column of map/claims.md:
`data/t04/x.npz` (Zenodo [10.5281/zenodo.1234567](https://doi.org/10.5281/zenodo.1234567), `t04.tar.gz`)
The release shown is the newest production release with a tar whose group path is the
artifact path, lies above it, or lies under it (an artifact that is a directory holding a
group). Sandbox releases are never shown: their DOIs (10.5072/...) do not resolve. Release
entries written before the tool recorded groups are matched with the current
zenodo.groups, or with the default one-group-per-data/ entry rule. The public pages that
opsci publish exports show the same release, in place of "(not published)".
Metadata¶
The title and creators come from CITATION.cff. Defaults: upload_type: dataset,
access_right: open, license: cc-by-4.0. The description names the project's public repo
and site when publish/manifest.yaml gives them (public_repo; site_url, else the GitHub
Pages URL of a github.com repo, as opsci publish derives it), and lists each tar with the
paths it holds and the task(s) they come from. With a public repo, the record also gets
related_identifiers: [{identifier: <repo URL>, relation: isSupplementTo, resource_type:
software}], which links the record back to the project. Override any field under zenodo.metadata in
data/MANIFEST.yaml. The project ships no .zenodo.json: Zenodo ignores CITATION.cff when
one exists, so metadata is sent through the API instead.
Tests¶
tests/test_zenodo.py runs against a local mock of the deposit API (tests/zenodo_mock.py).
The real sandbox test is marked manual:
tests/run_all --run-manual -k real_sandbox
It needs ~/.config/opsci/zenodo-sandbox.token.