Short answer: install marimo 0.24 or newer with huggingface_hub and Polars, create an HfApi object in a notebook cell, then open marimo’s Remote Storage panel to browse Hub datasets. Copy an hf://datasets/... path and pass it to pl.read_csv(), pl.read_parquet() or a lazy pl.scan_parquet() query. Public datasets need no token; private and gated datasets need authorised Hugging Face credentials.
marimo 0.24.0 was released on 17 August 2026 UTC with first-class Hugging Face Hub remote storage. A notebook can now detect huggingface_hub.HfApi, browse models, datasets and Spaces in the Files panel, and use Hub paths with data tools that understand the hf:// protocol.
What marimo 0.24 adds and what it does not
The new integration is primarily a discovery and path-handling improvement. When marimo sees a live HfApi object, its Remote Storage browser can expose repositories and files without making you hunt through a web page for every path. The actual data read still belongs to Polars, pandas, DuckDB or another compatible library.
It does not turn an entire Hub repository into one table, bypass a dataset’s access controls, guarantee that every repository contains Parquet, or make a large eager read fit in memory. You still need to choose a real file, understand its format and schema, follow its licence and dataset card, and decide whether an eager or lazy query is appropriate.
This workflow is useful when you want a reactive, inspectable alternative to a traditional notebook. If you are still deciding what kind of AI or data project to build, start with the site’s AI guides archive. For a model-training use case, the older Llama 2 fine-tuning guide provides context, although its model-specific commands should not be reused as current marimo instructions.
Prerequisites
- Python and a fresh virtual environment; the commands below use uv, which is the environment route shown in marimo’s current installation guide.
- marimo 0.24 or newer,
huggingface_huband a recent Polars release. Hugging Face documents nativehf://support in Polars 1.2 and newer. - Internet access to
huggingface.coand enough local memory or disk cache for the files you select. - A Hugging Face account and a least-privilege user access token only if the repository is private or gated. Gated data may also require you to accept its access conditions in the browser.
- Permission to use the dataset. Review its card, licence, personal-data notes, intended use and limitations before analysis or redistribution.
Do not put a production token into a notebook cell, commit it to Git or paste it into a screenshot. For a first run, use a small public file so authentication and dataset size cannot hide a basic environment problem.
1. Create an isolated marimo project
Open Terminal or PowerShell in a directory where you keep development projects, then create a dedicated environment:
mkdir marimo-hf-demo
cd marimo-hf-demo
uv init
uv add "marimo==0.24.0" "polars>=1.2" huggingface_hub
uv run marimo edit hf_dataset.pyThe marimo pin keeps this guide on its documented 0.24.0 release; the Polars minimum enables native Hub file access. Review release notes before deliberately updating the marimo pin. Let uv record resolved versions in its lockfile so another person can reproduce the environment. If you prefer marimo’s broader data-science bundle, its installation guide also documents marimo[recommended]; this focused setup installs only the packages used here.
2. Let marimo discover the Hugging Face Hub
Add this import cell to the new notebook:
import marimo as mo
import polars as pl
from huggingface_hub import HfApiIn the next cell, create and return an API client:
hf = HfApi()
hfRun both cells. Open the Files panel in marimo and select Remote Storage. According to marimo’s remote-storage guide, the editor automatically detects the HfApi object and lets you browse Hugging Face resources. Navigate to a dataset file and use the browser’s copy action rather than typing a long path from memory.
An hf:// dataset URI follows this shape:
hf://datasets/NAMESPACE/REPOSITORY/PATH/TO/FILEThe resource type is plural: datasets, models or spaces. Dataset names are not interchangeable with file paths, and path case matters. Hugging Face’s URI reference also supports an optional revision after the repository name, such as @main, a tag or a commit hash.
3. Read a small public CSV
marimo’s documentation uses the public Fish dataset as a compact example. Put the path and the read in separate reactive cells so changing the URI invalidates only downstream work:
DATA_URL = "hf://datasets/scikit-learn/Fish/Fish.csv"fish = pl.read_csv(DATA_URL)
fish.head()This is an eager read: Polars obtains the file before returning the DataFrame. That is reasonable for a small demonstration file. It is not a sensible default for a multi-gigabyte dataset.
4. Validate the data before analysing it
Do not treat a rendered table as proof that the right file arrived. Add small, explicit checks in separate cells:
fish.shapefish.schemafish.null_count()Compare the observed columns, types and row count with the repository’s dataset card or data viewer. A successful HTTP read can still deliver an unexpected revision, split or generated schema. For a durable project, record the repository name, exact file, revision, access date, licence and the validation result beside your analysis.
5. Use lazy Parquet reads for larger datasets
Hugging Face recommends Polars’ lazy API for larger Parquet data because it can push column selection and filters into the scan. The following pattern uses a file present in the official Stanford NLP IMDB repository. Its train shard was listed at about 21 MB during the source review; this example is a small projected aggregation, not a whole-repository download:
PARQUET_URL = (
"hf://datasets/stanfordnlp/imdb/"
"plain_text/train-00000-of-00001.parquet"
)
label_counts = (
pl.scan_parquet(PARQUET_URL)
.select("label")
.group_by("label")
.agg(pl.len().alias("rows"))
.sort("label")
.collect()
)
label_countsscan_parquet() builds a query plan; collect() executes it. Selecting only label before the aggregation gives Polars an opportunity to avoid materialising unrelated columns. If a repository contains many shards, Hugging Face also documents glob patterns, but inspect the matched file set before collecting it so a wildcard does not unexpectedly span every split.
Many Hub datasets have an automatically converted Parquet branch. Hugging Face identifies that revision as ~parquet, which can appear in a URI after @. Treat it as a generated representation, verify its configuration and split paths, and pin a stable revision when reproducibility matters.
6. Authenticate for private or gated data
For a private repository, create a read-only user token in Hugging Face settings and authenticate outside the notebook:
uv run hf auth loginPaste the token only into the CLI prompt. The Hub client and compatible filesystem integrations can use the cached credential without exposing it in the Python source. Hugging Face documents credential priority and also recognises the HF_TOKEN environment variable, but a notebook cell that prints its environment can leak that value. Use the narrowest token scope possible and revoke the token when the project ends.
A token cannot grant access you have not received. For a gated dataset, sign in to the dataset page, review its conditions and request or accept access first. A 401 generally points to missing or invalid authentication; a 403 can mean the account lacks repository or gating permission.
7. Pin a revision for reproducible work
A path without a revision usually follows the repository’s default branch, which can change after you publish results. Once an exploratory query works, copy a commit identifier from the Hub and add it after the repository name:
hf://datasets/NAMESPACE/REPOSITORY@COMMIT_SHA/PATH/TO/FILEKeep the resolved Python packages in uv.lock as well. A commit-pinned dataset plus a locked environment does not guarantee identical results across every platform, but it removes two major sources of silent drift.
8. Run a clean validation pass
- Save
hf_dataset.pyand restart marimo’s kernel. - Run the notebook from the first import cell in dependency order.
- Confirm Remote Storage appears after the
HfApicell executes. - Confirm the URI, schema, shape and null checks match the selected dataset card.
- Run the lazy Parquet aggregation and watch memory use before attempting a broader scan.
- Open a read-only app view with
uv run marimo run hf_dataset.pyand confirm the outputs render without editor-only state.
For any published analysis, add assertions appropriate to the dataset: expected columns, allowed label values, uniqueness of keys, plausible date ranges and an acceptable missing-data threshold. The three generic checks above verify transport and basic shape; they do not establish analytical quality.
Troubleshooting
| Problem | Likely cause | What to check |
|---|---|---|
| Remote Storage does not show Hugging Face | Older marimo or the API object has not run | Check uv run marimo --version, update to 0.24 or newer, execute the HfApi() cell and reopen the Files panel. |
| “Unknown protocol” or filesystem error | Old Polars or malformed URI | Use Polars 1.2 or newer and the plural hf://datasets/NAMESPACE/REPO/path form. |
| File not found | Wrong revision, split, case or physical filename | Browse the repository in Remote Storage and copy the exact file URI. Do not assume a dataset ID is itself a file. |
| 401 or 403 | Missing token, insufficient scope or unaccepted gate | Run uv run hf auth whoami, confirm read access and accept the dataset’s access conditions. |
| Parquet schema or Arrow error | Wrong format, dependency mismatch or inconsistent shards | Confirm the extension and repository card, update the locked environment deliberately, and test one shard before a glob. |
| Read is slow or memory use spikes | Eager read or overly broad scan | Use scan_parquet(), select columns, filter early and inspect the query before collect(). |
| Notebook works only after manual clicks | Hidden state or cells run out of order | Restart the kernel, run from the top and make every input an explicit reactive dependency. |
Rollback and credential cleanup
Stop the marimo process before cleanup. If you no longer need Hub access on the machine, run uv run hf auth logout and revoke the project token in your Hugging Face account settings. If you temporarily used an environment variable, remove it in the same shell with unset HF_TOKEN on macOS or Linux, or Remove-Item Env:HF_TOKEN in PowerShell.
Remove or disable the HfApi cell if you do not want the notebook to expose Hub browsing. Keep hf_dataset.py, pyproject.toml and uv.lock together if the work needs to be reproduced. If the project is disposable, move its directory to the operating system’s Trash or Recycle Bin only after exporting any results you need. Authentication cleanup does not automatically delete Hugging Face cache files; inspect cache locations and sizes before removing anything shared with other projects.
Frequently asked questions
Does hf:// download the whole dataset?
Not necessarily. The URI names a repository file or file pattern. An eager reader such as read_csv() fetches the selected file, while a lazy Parquet scan can request only needed columns and row groups when the backend supports it. A broad glob can still touch many shards.
Do I need the Hugging Face Datasets library?
No, not for the direct Polars workflow in this guide. huggingface_hub supplies the API connection and Polars reads the file. Install the separate datasets package when you need its builders, streaming abstractions or transforms.
Can marimo access private Hugging Face datasets?
Yes, if the signed-in account has access and the client can find a valid token. Private and gated are different: a gated dataset can require a separate approval or terms acceptance even when authentication works.
Can I write data back to the Hub?
Hugging Face supports uploads through its APIs and compatible tools, but this guide is deliberately read-only. Publishing a dataset requires repository permissions, a data card, licence decisions, privacy review and a deliberate commit workflow.
Is a marimo notebook a normal Python file?
Yes. marimo stores notebooks as Python files and builds a dependency graph between cells. Keep secrets outside the file and use the clean-run test to catch hidden state before sharing it.
Bottom line
marimo 0.24 makes Hugging Face datasets easier to discover, but a reliable workflow still depends on exact file paths, least-privilege authentication, lazy reads for large Parquet data and explicit validation. Start with a small public CSV, confirm the Remote Storage connection, graduate to a projected lazy scan, then pin both the dataset revision and Python environment before relying on the result.
Primary project sources
- marimo 0.24.0 release notes
- marimo remote-storage guide
- marimo installation guide
- marimo quickstart and application mode
- Hugging Face Hub: use Polars with datasets
- Hugging Face Hub: authenticate Polars access
- Hugging Face Hub URI reference
Sources reviewed 10 September 2026 against the official documentation linked above. Record versions and validate the workflow in your own environment.