ARC (file format)
View on WikipediaThis article is missing information about the file format specifications and its byte signature. (February 2019) |
| ARC | |
|---|---|
| Filename extension |
.arc, .ark |
| Internet media type |
application/octet-stream |
| Uniform Type Identifier (UTI) | public.archive.arc |
| Developed by | System Enhancement Associates |
| Type of format | Data compression |
ARC is a lossless data compression and archival format by System Enhancement Associates (SEA). The file format and the program were both called ARC. The format is known as the subject of controversy in the 1980s, part of important debates over what would later be known as open formats.
ARC was extremely popular during the early days of the dial-up BBS. ARC was convenient as it combined the functions of the SQ program to compress files and the LU program to create .LBR archives of multiple files. The format was later replaced by the ZIP format, which offered better compression ratios and the ability to retain directory structures through the compression/decompression process.
The .arc filename extension is often used for several unrelated file archive-like file types. For example, the Internet Archive used its own ARC format to store multiple web resources into a single file.[1][2] The FreeArc archiver also uses a .arc extension, but uses a completely different file format. Nintendo uses an unrelated "ARC" format for resources, such as audio, or text, in GameCube and Wii games. Several unofficial extractors exist for this type of ARC file.[which?][citation needed]
History
[edit]In 1985, Thom Henderson of System Enhancement Associates wrote a program called ARC,[3] based on earlier programs such as ar, that not only grouped files into a single archive file but also compressed them to save disk space, a feature of great importance on early personal computers, where space was very limited and modem transmission speeds were very slow. The archive files produced by ARC had file names ending in ".ARC" and were thus sometimes called "arc files".
The source code for ARC was released by SEA in 1986 and subsequently ported to Unix and Atari ST in 1987 by Howard Chu. This more portable codebase was subsequently ported to other platforms, including VAX/VMS and IBM System/370 mainframes. Howard's work was also the first to disprove the prevalent belief that Lempel-Ziv encoded files could not be further compressed. Additional compression could be achieved by using Huffman coding on the LZW data, and Howard's version of ARC was the first program to demonstrate this property. This hybrid technique was later used in several other compression schemes by Phil Katz and others.
Later, Phil Katz developed his own shareware utilities, PKARC and PKXARC, to create archive files and extract their contents. These files worked with the archive file format used by ARC and were significantly faster than ARC on the IBM-PC platform due to selective assembly-language coding. Unlike SEA, which combined archive creation and archive file extraction in a single program, Katz divided these functions between two separate utilities, reducing the amount of memory needed to run them. PKARC also allowed the creation of self-extracting archives, which could unpack themselves without requiring an external file extraction utility.
Following the System Enhancement Associates, Inc. vs PKWARE Inc. and Phillip W. Katz lawsuit, SEA withdrew from the shareware market and developed ARC+Plus.[4] This version included a full-screen user interface, with the last known version being 7.12.[5] SEA was eventually sold to an unspecified Japanese company in 1992.[6]
The ARC format is no longer common on PC desktops, but most antivirus scanners can still uncompress any ARC archives found in order to detect viruses within the compressed files.
Lawsuits
[edit]In the late 1980s a dispute arose between SEA, maker of the ARC program, and PKWARE, Inc. (Phil Katz Software). SEA sued Katz for trademark and copyright infringement. An independent software expert, John Navas, was appointed by the court to compare the two programs, and stated that PKARC was a derivative work of ARC, pointing out that comments in both programs were often identical, including spelling errors.[7]
On August 2, 1988, the plaintiff and defendants announced a settlement of the lawsuit, which included a Confidential Cross-License Agreement under which SEA licensed PKWARE for all the ARC-compatible programs published by PKWARE during the period beginning with the first release of PKXARC in late 1985 through July 31, 1988, in return for $62,500, which at the time was an undisclosed payment amount. In the agreement, PKWARE paid SEA to obtain a license that allowed the distribution of PKWARE's ARC-compatible programs until January 31, 1989, after which PKWARE would not license, publish or distribute any ARC compatible programs or utilities that process ARC compatible files. In exchange, PKWARE licensed SEA to use its source code for PKWARE's ARC-compatible programs. PKWARE also agreed to cease any use of SEA's trademark "ARC" and to change the names or marks used with PKWARE's programs to non-confusing designations. The remaining details of the agreement were sealed. In reaching the settlement, the defendants did not admit any fault or wrongdoing.[8] The Wisconsin court order showed defendants were ordered to pay damages to plaintiff for defendants' acts of infringing Plaintiff's copyrights, trademark, and acts of unfair trade practices and unfair competition.[9]
The leaked agreement document revealed under the settlement terms, the defendants had paid plaintiff $22,500 for past royalty payments, and $40,000 for expense reimbursements. In addition, defendants would pay plaintiff a royalty fee of 6.5% of all revenue received for ARC compatible programs on all orders received after the effective date of this Agreement, such revenue including any license fees or shareware registrations received after the expiration of the license, for ARC compatible programs. In exchange, plaintiff would also pay a commission in the amount of 6.5% of any license fees received by plaintiff from any licensee referred to plaintiff by defendants, whether before or after the license termination date.[10]
After the lawsuit, PKWARE released one last version of his PKARC and PKXARC utilities under the new names "PKPAK" and "PKUNPAK", and from then on concentrated on developing the separate programs PKZIP and PKUNZIP, which were based on new and different file compression techniques and archive file formats. However, following the renaming, SEA filed a lawsuit against PKWARE for contempt, for continually using plaintiff's protected mark ARC, by turning ARC from noun into verb in the PKPAK manual.[11] The U.S. district court of the East District of Wisconsin ruled SEA's motion was denied, and the defendant was entitled to recover the legal cost of $500.[12]
The SEA vs. PKWARE dispute quickly expanded into one of the largest controversies the BBS world ever saw.[13] The suit by SEA angered many shareware users who perceived that SEA was a "large, faceless corporation" and Katz was "the little guy". In fact, at the time, both SEA and PKWARE were small home-based companies. However, the community largely sided with Katz, due to the fact that SEA was attempting to retroactively declare the ARC file format to be closed and proprietary. Katz received positive publicity by releasing the APPNOTE.TXT specification documenting the ZIP file format, and declaring that the ZIP file format would always be free for competing software to implement. The net result was that the ARC format quickly dropped out of common use as the predominant compression format that PC-BBSs used for their file archives, and after a brief period of competing formats, the ZIP format was adopted as the predominant standard.
In an interview, Thom Henderson of SEA said that the main reason he dropped out of software development was because of his inability to emotionally cope with what he claimed was the hate-mail campaign launched against him by Katz.[14]
See also
[edit]References
[edit]- ^ "13. Internet Archive ARC files". Retrieved 2012-07-17.
- ^ "Internet Archive: ARC File Format Reference". Retrieved 2012-07-17.
- ^ "Phil Katz". www.esva.net. Retrieved 15 March 2018.
- ^ Vaughan-Nichols, Steven J. (1 November 1991). "ARC+Plus 7.12. (Software Review) (one of seven evaluations of data compression utility programs in 'Space Savers: Data Compression Utilities') (Evaluation)". Computer Shopper (US magazine). Archived from the original on 4 November 2012. Retrieved 15 March 2018.
- ^ "Compression packages (results and site)". www.bio.net. Retrieved 15 March 2018.
- ^ "Thom Henderson". www.esva.net. Retrieved 2018-10-16.
- ^ Response, Fredric L. Rice, Organized Crime Civilian. "Thom Henderson, president System Enhancement Associates voice: (201) 473-5153 data: (201)". www.skepticfiles.org. Archived from the original on 30 June 2014. Retrieved 15 March 2018.
{{cite web}}: CS1 maint: multiple names: authors list (link) - ^ "Joint press release". Retrieved 15 March 2018.
- ^ System Enhancement Associates, Inc. v. PKWare, Inc. and Phillip W. Katz, No. 88-C-447, Judgment for Plaintiff on Consent, E.D. Wisc. (Aug. 1., 1988)
- ^ "System Enhancement Associates vs. PKware, Inc Confidential Cross-License Agreement". Retrieved 15 March 2018.
- ^ "System Enhancement Associates vs. PKware, Inc". Retrieved 15 March 2018.
- ^ "United States District Court Eastern District of Wisconsin Case No. 88-C-447". Retrieved 15 March 2018.
- ^ BBS Documentary, Episode 8, [1], Accessed as of 13.07.2012
- ^ BBS: The Documentary, Episode 3.03 Compression.
External links
[edit]- ARC file format description
- ARC — free software Linux/Unix port of the .arc compression program
- nomarch — another free software .arc compression program for Linux/Unix
- The BBS Documentary: Compression on YouTube — A documentary by Jason Scott that discusses ARC history, in the context of BBS, with notes: "Controversy: Lawsuits: SEA vs. PKWARE"
- UnARC — ARC unpacker for Atari 8-bit computers and Atari DOS
ARC (file format)
View on GrokipediaHistory
Origins and Development
The ARC file format originated in 1985, when Thom Henderson, founder of System Enhancement Associates (SEA), developed the ARC utility program for MS-DOS systems.[5][1] This shareware tool combined file archiving with data compression, addressing limitations of earlier standalone utilities like SQ (Squeeze) for packing and basic archivers for grouping files.[7] ARC's initial implementation employed Lempel–Ziv–Welch (LZW) compression, an adaptive dictionary-based algorithm that improved efficiency over prior methods, making it suitable for distributing software via limited-bandwidth dial-up bulletin board systems (BBS).[7][2] SEA marketed ARC as a commercial product while distributing it under shareware terms, requiring users to register for full features after a trial period.[5] The format quickly gained prominence in the BBS community from 1985 to 1989, becoming the dominant standard for compressing and archiving files due to its balance of compression ratios and extraction speed on era hardware.[3] Early versions supported multiple packing methods, including a "trimmed" Huffman coding variant for enhanced performance on repetitive data.[8] Development progressed through iterative releases, with SEA issuing updates that refined the file header structure for better metadata handling, such as file dates, sizes, and CRC checksums for integrity verification.[1] By late 1989, version 7.x introduced optimized compression schemes, though the core format remained backward-compatible to maintain interoperability across SEA's ecosystem.[8] These enhancements solidified ARC's role as a precursor to later formats, influencing subsequent archivers amid growing demand for efficient file distribution in pre-internet computing.[2]Evolution and Updates
The ARC utility, initially released in 1985 by System Enhancement Associates (SEA), saw iterative updates that primarily enhanced compression efficiency through new algorithms while maintaining core file structure compatibility. Early versions supported basic methods such as "packing" for uncompressed or lightly reduced storage and "squeezing," an LZW variant for dictionary-based compression. By version 6.02, released on October 3, 1988, the program included refinements to existing schemes like "crunching," which combined statistical modeling with Huffman coding for better ratios on certain data types.[9] Version 6.00, documented in January 1989, formalized these capabilities in an updated manual, emphasizing multi-file archiving with optional password protection and self-extracting variants.[10] The final major release, version 7.x (e.g., ARC712.EXE), emerged in late 1989 or early 1990 as part of the ARC+Plus package, introducing "Trimmed" (method ID 10) as the default compression. This scheme integrated RLE90 run-length encoding to preprocess repeats, a 4KB LZ77 history buffer for match distances up to 4096, and adaptive Huffman coding on 314 symbols (including literal bytes and length codes from 3 to 59), yielding superior performance over prior methods without altering the overarching header format.[8] Subsequent development ceased amid competition from PKZIP, with no format-altering updates beyond v7; later efforts like SuperARC variants for platforms such as Atari 8-bit were derivatives based on v5.0 code rather than official SEA evolutions.[9] The format's flexibility in per-file method selection (via one-letter codes in headers) allowed backward compatibility, but its evolution prioritized algorithmic gains over structural overhauls.[8]Technical Specifications
File Structure and Headers
The ARC file format employs a sequential structure without a central directory, consisting of one or more archive entries (members) followed by an end-of-archive marker comprising the bytes0x1A 0x00.[11] Each entry begins with a fixed 26-byte header that encapsulates metadata for the stored file, including its compression details and attributes.[12] The absence of a global header or index requires extraction tools to parse entries linearly from the file's start.
The header commences with the signature byte 0x1A at offset 0x00, serving as the entry delimiter, immediately followed at offset 0x01 by a single byte denoting the compression method—common values include 0x03 for packing (run-length encoding), 0x04 for squeezing (RLE combined with Huffman coding), and 0x08 for crunching (RLE with LZW dictionary compression up to 12-bit codes).[4][12] Offsets 0x02 through 0x0D hold the filename as 12 ASCII characters, typically padded with nulls or spaces if shorter, accommodating MS-DOS 8.3 naming conventions with an implicit extension period.[12] The compressed data length follows as a 4-byte little-endian unsigned integer at offsets 0x0E to 0x11, succeeded by a 4-byte MS-DOS timestamp (dword) at 0x12 to 0x15—encoding year (offset from 1980), month, day, hour, minute, and seconds/2 in bit-packed format. A 16-bit CRC checksum of the compressed data occupies offsets 0x16 to 0x17 (using polynomial 0x1021 or equivalent for validation), and the original uncompressed file size concludes the header as another 4-byte little-endian value at 0x18 to 0x1B.[4][12]
| Offset (hex) | Size (bytes) | Field | Description |
|---|---|---|---|
| 00 | 1 | Signature | Fixed value 0x1A marking entry start.[12] |
| 01 | 1 | Compression method | Byte ID for algorithm (e.g., 0x08 for LZW-based crunching).[4][12] |
| 02–0D | 12 | Filename | ASCII name, padded if necessary.[12] |
| 0E–11 | 4 | Compressed size | Little-endian length of following data.[12] |
| 12–15 | 4 | MS-DOS timestamp | Bit-packed date/time (year-1980:7 bits, month:4, day:5, hour:5, min:6, sec/2:5).[12] |
| 16–17 | 2 | CRC-16 | Checksum of compressed data (polynomial 0x1021).[4] |
| 18–1B | 4 | Original size | Little-endian uncompressed file length.[12] |
Compression Algorithms
The ARC file format supports multiple lossless compression algorithms, applied independently to each archived file to minimize overall size, with the method selected based on empirical testing for the best ratio. The compression type is specified by an 8-bit identifier in the file header, allowing flexibility across different data types. Early versions of the ARC utility, released by System Enhancement Associates starting in 1985, relied on simpler techniques like run-length encoding and Huffman coding, while later iterations from 1989 onward incorporated dictionary-based methods inspired by LZ77 variants.[4][8] One foundational method, known as "Packing" (compression type 3), uses run-length encoding (RLE) with 0x90 as a special marker: a literal 0x90 is encoded as 0x90 followed by 0x00, while repeats of the prior byte are encoded as 0x90 followed by the count (up to 255 additional occurrences). This pre-processing step reduces redundancy in repetitive data but offers limited compression for diverse content.[4] "Squeezing" (type 4) builds on RLE by applying static Huffman coding to the encoded stream, where the code tree is prefixed with a 16-bit node count followed by an array of 16-bit node indices (negative values indicating leaf symbols). Decoding proceeds bit-by-bit in least-significant-bit order, making it suitable for files with skewed symbol distributions post-RLE.[4] "Crunched" (type 8) extends RLE with Lempel-Ziv-Welch (LZW) dictionary compression using an 8-kilobyte buffer, starting with 9-bit codes that dynamically expand to 12 bits; a reset code (0x0100) clears the dictionary when full, with bit widths adjusting accordingly. This approach, akin to early adaptive dictionary schemes, improved ratios for textual and structured data but was computationally intensive for the era's hardware.[4][13] Advanced methods in extensions like PAK (version 2.0, July 1989) include "Distilled," which integrates LZ77 sliding-window matching (8KB buffer, up to 60-byte matches, supporting prehistory offsets) with dual Huffman codebooks: a dynamic one for literals (0-255), match lengths (3-60), and stop codes, plus a fixed one for offsets derived from decompressed byte contexts. The stream is bit-packed, with offsets computed variably (0-7 extra bits based on recent data).[14] "Trimmed," default in ARC version 7 (late 1989), preprocesses with RLE90 before LZ77 (4KB buffer, matches 3-59 bytes) encoded via adaptive Huffman trees over 314 symbols (literals plus length codes, with 256 as stop). This hybrid yielded competitive ratios, outperforming simpler modes on mixed workloads while remaining decode-efficient.[8]Software and Compatibility
Original SEA ARC Program
The original SEA ARC program was developed by Thom Henderson and released in 1985 by System Enhancement Associates (SEA) as shareware software primarily for MS-DOS systems on IBM PCs.[15][5] Written in C using the C86 compiler, it enabled users to bundle multiple files into a single .arc archive while applying lossless compression to reduce file sizes, facilitating efficient storage and transfer over dial-up connections.[5] The program quickly gained traction among operators of electronic bulletin board systems (BBSes), which dominated pre-Internet file sharing, as it significantly decreased download times for software distributions—for instance, version 5.12 compressed an MS-DOS 3.20 installation archive by over 20%.[15][5] Key features included command-line operations such as adding files (a), extracting (e), listing contents (l), and deleting entries (d), with support for multiple compression methods indicated by a method byte in the file header, including storage (method 0), packing (method 1), and squeezing (method 2) for varying data types.[1] Later updates introduced additional algorithms like crunching (method 3) and squashing (method 4), alongside optional password protection for archives, though forgotten passwords could render files irretrievable without recovery tools.[5][1] SEA released the program's source code in 1986, promoting ports to platforms like Unix and Atari ST by developers such as Howard Chu in 1987, which extended its utility beyond DOS environments.[15][1]
By 1989, ARC had evolved to version 6.00, incorporating refinements for better compatibility and performance, but the core program remained a staple for archival tasks until displaced by faster alternatives.[10] Its design emphasized simplicity and effectiveness for resource-constrained 1980s hardware, supporting features like file commenting and multi-volume archives, though it lacked native encryption beyond basic passwords.[1] The program's shareware model, requiring registration for full access, contributed to its widespread adoption while generating revenue for SEA.[5]
Third-Party Implementations
PKARC, developed by Phil Katz in 1987, was a prominent third-party implementation designed for compatibility with the ARC format, offering faster compression and extraction speeds than the original SEA software.[16] It supported the same file structure and compression methods, allowing users to create and manipulate .ARC archives, but its release prompted a copyright infringement lawsuit from SEA alleging unauthorized code duplication.[17] Other contemporary third-party tools, such as QUARK and SQUASH, also utilized the ARC format for archiving and compression on MS-DOS systems, extending compatibility without direct affiliation to SEA.[17] These implementations maintained the core header and packing mechanisms of ARC, enabling interoperability in the BBS and shareware ecosystems of the late 1980s. In modern contexts, compatibility is primarily through extraction tools rather than full read-write support, as the format's proprietary elements limited widespread adoption post-lawsuit. IZArc, a free Windows utility, can decompress legacy ARC files alongside numerous other formats.[18] Command-line tools like unarc, available on Linux distributions, provide extraction capabilities for older ARC archives by parsing the 1A header and file entries.[19] These tools focus on read access for archival preservation, reflecting the format's obsolescence for new creations since the early 1990s.[20]Modern Tools and Emulation
IZArc, a free Windows-based compression utility, supports extraction and creation of files in the legacy SEA ARC format, allowing users to handle these archives on modern operating systems without requiring DOS emulation.[18] Other tools, such as certain Linux command-line utilities like thearc package, provide compatibility for unpacking older ARC files, primarily for archival and retro computing purposes.[21]
For scenarios where native support is absent or incomplete, DOS emulators enable execution of the original SEA ARC program (ARC.EXE) from the 1980s. DOSBox, a widely used x86 emulator, runs the utility accurately on Windows, Linux, and macOS, facilitating direct extraction via commands like ARC e filename.arc in an emulated MS-DOS environment.[22] Enhanced variants like DOSBox-X offer improved compatibility for period-specific behaviors, such as handling compression algorithms like LZW or Huffman coding inherent to early ARC versions.[23]
Online platforms like PCjs provide browser-based emulation of IBM PC hardware, permitting users to upload and process ARC files interactively without local installation, preserving access to self-extracting archives and multi-volume sets.[5] These emulation approaches maintain fidelity to the format's original specifications, including variable-length headers and packing methods, though they may require sourcing authentic ARC executables from trusted retro archives to avoid compatibility issues with modified versions.[5]