Apache Parquet file format
.parquet · columnar data file · introduced 2013 by Twitter & Cloudera (now Apache)
What is a Apache Parquet file?
Parquet is the dominant columnar data file format for analytics. Each file groups rows into row-groups and stores each column as its own compressed page, allowing query engines to read only the columns they need. DuckDB, Spark, BigQuery, Snowflake and pandas all read and write Parquet natively.
Most readers reach this page after double-clicking a .parquet file and finding their operating system unsure what to do with it. The good news: Apache Parquet is a well-understood format with mature tooling on every major platform. It is supported in one click by mainstream archivers on Windows, macOS, and Linux.
How to extract a Apache Parquet file
Opening a .parquet file is straightforward on every modern operating system. The exact steps depend on which platform you are on:
On Windows
- Locate the
.parquetfile in File Explorer. - Right-click the file. If you have 7-Zip installed, choose 7-Zip → Extract Here. On Windows 11 the built-in extractor handles ZIP, 7z, RAR and several other formats directly through Extract All…
- If prompted, pick a destination folder and click Extract.
- The extracted files appear in a folder next to the original archive.
On macOS
- Find the
.parquetfile in Finder. - Install The Unarchiver from the Mac App Store, or Keka for both creating and extracting.
- Right-click the file → Open With → The Unarchiver (or just double-click after Unarchiver is set as the default).
- The extracted folder appears next to the archive.
On Linux
- Use your file manager: right-click the archive → Extract Here (file-roller on GNOME, Ark on KDE).
- On the command line, run one of:
7z x archive.parquet # works for almost any format
- The files appear in the current directory.
How to create a Apache Parquet file
To pack files into a .parquet archive, you typically need an archiver that supports writing this format. Not every tool can create every format — for example, only WinRAR can create real RAR files. The sidebar lists the software known to read and write Apache Parquet.
On the command line, the canonical create command is:
# Use 7-Zip, PeaZip, or the format's reference tool to create .parquet files.
If you regularly create archives at scale, a dedicated archive automation suite can wrap this command in scheduled jobs, integrity checks, and offsite replication.
Strengths
- Columnar — read only the columns you need
- Multiple compression algorithms
- Schema and stats in the footer
Weaknesses and limitations
- Append unfriendly
- Many small files perform poorly
Typical use cases
- Data lake storage
- Analytics pipelines
- BigQuery / Athena / Spark workloads
Technical details
Apache Parquet uses the Snappy / GZIP / Zstandard / LZ4 algorithm and is identified on disk by the magic byte sequence 50 41 52 31. The standard MIME type is application/vnd.apache.parquet. The format is an open specification and is free for any use.
Frequently asked questions
Is Apache Parquet safe to open?
Archive files are containers — they are only as safe as what they contain. A .parquet from a trusted source is fine; one received from a stranger can hold malicious executables, scripts, or files engineered to exploit bugs in your archiver. Always keep your archiver updated, and never run unknown executables that come out of an archive. A safe workflow is to extract into a sandbox and scan the contents before opening anything.
What is the maximum file size?
Modern implementations support archives in the multi-terabyte range. The practical limit is your filesystem (FAT32 caps at 4 GB; ext4, NTFS, APFS, and exFAT all comfortably handle multi-TB files).
Can I password-protect a Apache Parquet?
Some archivers can wrap Apache Parquet files in an encrypted container, but the format itself does not always include built-in encryption. Check your archiver's documentation.
How does Apache Parquet compare to other formats?
See our format comparison pages such as ZIP vs 7z, ZIP vs RAR, and tar.gz vs tar.xz for side-by-side feature tables.