Parquet viewer

Open Parquet files and datasets locally. Inspect rows and schema, filter data, and export JSON or CSV.

What the viewer shows

Every part of the file is decoded in this tab: the column schema with logical types and repetition, the row-group count, and the compression codecs actually used. Row values are shown exactly as stored — 64-bit integers and decimals stay exact strings instead of losing precision to a JavaScript number, binary columns are shown as hexadecimal, timestamps keep their original unit, and nulls are shown as null rather than as an empty cell.

The filter box searches every column of every row, not just the page on screen. Rows are paginated 50 at a time so a large file stays responsive, and the JSON and CSV download buttons write out every matching row — all of them when no filter is active, and exactly the filtered set when one is.

Files, parts and partitioned folders

A Parquet table is often many files. Select several .parquet parts, or open a whole folder, and the viewer validates every part and assembles them into one table — with the included files, their row counts and their sizes listed so you can see what was read. Non-Parquet files found in a folder, such as _SUCCESS or .crc marker files, are listed explicitly as excluded rather than quietly ignored.

Hive-style folder segments such as year=2026/month=09 become real columns when the files themselves do not carry them. Percent escapes are decoded, leading zeroes are preserved as strings, and __HIVE_DEFAULT_PARTITION__ reads as null. If any part is corrupt, its schema is incompatible, or a partition value conflicts, the whole selection is rejected with the offending path named. You never get a silently truncated table.

Limits

The viewer holds one dataset at a time. Up to 25 MB of selected files (1 MB = 1,000,000 bytes). Parquet limits also apply to expanded data.

Per file or dataset: up to 256 files, 25 MB of selected files in the browser, 256 MiB uncompressed data, 500,000 rows, and 5 million cells. All processing stays local. Rows are decoded and held in memory, so these ceilings are checked before anything is read: a selection over the limit is refused outright instead of being cut down to fit.

Nothing is uploaded

There is no server side to this page. The file never leaves your device: it is decoded by a Web Worker in this tab, and the Content-Security-Policy on this site allows connections to no origin but its own, so no script here could send your data anywhere even if it tried. You can check that in DevTools, or simply turn your wifi off and open a file anyway — the page is cached for offline use after your first visit.

Questions

Does opening a Parquet file here upload it?

No. The file is read and decoded entirely in your browser tab by a Web Worker. No part of it is sent to a server, and the site loads no third-party scripts.

Can it open a folder of Parquet parts?

Yes. Choose Open Parquet folder, or drop a directory onto the page. Subfolders are discovered recursively, every part is validated before decoding, and Hive-style partition folders such as year=2026 become columns.

How large a file can it open?

Up to 25 MB of selected files, 256 parts, 256 MiB of uncompressed data, 500,000 rows and 5 million cells. Larger selections are refused with a clear message rather than partly loaded.

Are big integers and decimals shown exactly?

Yes. INT64 and DECIMAL values are kept as exact strings, binary values are shown as hexadecimal, and timestamps retain their stored precision, so nothing is rounded on the way to the screen.

Can I compare two Parquet datasets?

Yes, on the main cruxdiff comparison tool. It matches records by a key column you choose, reports added, removed and modified rows, and treats file names, part counts and row order as noise rather than changes.

Compare two Parquet datasets · All comparison tools · How the privacy claim is enforced