← All formats

Apache ORC file format

.orc  ·  columnar data file  ·  introduced 2013 by Hortonworks

What is a Apache ORC file?

ORC (Optimized Row Columnar) is a Hadoop-ecosystem columnar format from Hortonworks, designed for Hive workloads. Like Parquet, it stripes data column-by-column with rich indexing and statistics, but with a different row-group layout that often performs better on Hive itself. Outside the Hive world, Parquet is more common.

Most readers reach this page after double-clicking a .orc file and finding their operating system unsure what to do with it. The good news: Apache ORC is a well-understood format with mature tooling on every major platform. It is supported in one click by mainstream archivers on Windows, macOS, and Linux.

How to extract a Apache ORC file

Opening a .orc file is straightforward on every modern operating system. The exact steps depend on which platform you are on:

On Windows

  1. Locate the .orc file in File Explorer.
  2. Right-click the file. If you have 7-Zip installed, choose 7-Zip → Extract Here. On Windows 11 the built-in extractor handles ZIP, 7z, RAR and several other formats directly through Extract All…
  3. If prompted, pick a destination folder and click Extract.
  4. The extracted files appear in a folder next to the original archive.

On macOS

  1. Find the .orc file in Finder.
  2. Install The Unarchiver from the Mac App Store, or Keka for both creating and extracting.
  3. Right-click the file → Open WithThe Unarchiver (or just double-click after Unarchiver is set as the default).
  4. The extracted folder appears next to the archive.

On Linux

  1. Use your file manager: right-click the archive → Extract Here (file-roller on GNOME, Ark on KDE).
  2. On the command line, run one of:
    7z x archive.orc    # works for almost any format
  3. The files appear in the current directory.

How to create a Apache ORC file

To pack files into a .orc archive, you typically need an archiver that supports writing this format. Not every tool can create every format — for example, only WinRAR can create real RAR files. The sidebar lists the software known to read and write Apache ORC.

On the command line, the canonical create command is:

# Use 7-Zip, PeaZip, or the format's reference tool to create .orc files.

If you regularly create archives at scale, a dedicated archive automation suite can wrap this command in scheduled jobs, integrity checks, and offsite replication.

Strengths

  • Strong predicate pushdown
  • Multiple compressors
  • Rich indexes

Weaknesses and limitations

  • Less popular than Parquet outside Hive

Typical use cases

  • Hive data warehouses
  • Trino / Presto queries on HDFS
  • Hadoop ETL

Technical details

Apache ORC uses the Zstandard / Snappy / zlib / LZ4 algorithm and is identified on disk by the magic byte sequence 4F 52 43. The standard MIME type is application/x-orc. The format is an open specification and is free for any use.

Frequently asked questions

Is Apache ORC safe to open?

Archive files are containers — they are only as safe as what they contain. A .orc from a trusted source is fine; one received from a stranger can hold malicious executables, scripts, or files engineered to exploit bugs in your archiver. Always keep your archiver updated, and never run unknown executables that come out of an archive. A safe workflow is to extract into a sandbox and scan the contents before opening anything.

What is the maximum file size?

Modern implementations support archives in the multi-terabyte range. The practical limit is your filesystem (FAT32 caps at 4 GB; ext4, NTFS, APFS, and exFAT all comfortably handle multi-TB files).

Can I password-protect a Apache ORC?

Some archivers can wrap Apache ORC files in an encrypted container, but the format itself does not always include built-in encryption. Check your archiver's documentation.

How does Apache ORC compare to other formats?

See our format comparison pages such as ZIP vs 7z, ZIP vs RAR, and tar.gz vs tar.xz for side-by-side feature tables.

Related formats