Convert webarchive to html

Convert WEBARCHIVE to HTML

Extract Safari WEBARCHIVE content as editable, browser-compatible HTML.

Make HTML files online

We can't read WEBARCHIVE files yet, so this conversion isn't available. If you can export your work to one of these formats - or others - we'll turn it into HTML:

How to convert webarchive to html file

A Safari archive may need conversion when an editor, web server, or non-Apple browser requires an ordinary .html file. Converting it also makes the document easier to inspect, edit, or share without Safari.

What the .webarchive format is

A .webarchive file is usually a binary Apple property-list archive created by Safari or another WebKit-based application. It can contain the page HTML, images, stylesheets, scripts, fonts, and metadata in one file, allowing an offline copy to be reopened with much of its original appearance.

Safari creates one through File → Save As when Web Archive is selected as the format. WEBARCHIVE is an Apple-associated container, not a standard format supported by all browsers or web servers.

What the .html format is

HTML is the standard markup used for web pages. An .html file can be opened in modern browsers, edited in a code editor, published to a web server, and processed by indexing or document tools without Safari.

HTML normally contains references to separate images, CSS, JavaScript, and font files. Changing the extension or extracting only the main resource does not automatically create a complete offline copy.

Convert it with Safari on macOS

  1. Open the .webarchive file in Safari.
  2. Select File → Save As.
  3. In the format menu, choose Page Source or Page Source HTML, depending on the Safari version.
  4. Save the file with an .html extension.

This is the quickest desktop method, but Safari versions differ. Some installations expose only Web Archive in the save dialog; in that case, use the extraction method below. The exported source may not include every archived subresource in a separate local directory.

Extract the main resource with Python

Python’s standard library can read the binary property-list structure without installing a converter. Save this as extract_webarchive.py:

import plistlib
import sys

source, destination = sys.argv[1], sys.argv[2]
with open(source, 'rb') as file:
    archive = plistlib.load(file)

main = archive['WebMainResource']
data = main['WebResourceData']
if not isinstance(data, bytes):
    raise TypeError('WebResourceData is not binary data')

with open(destination, 'wb') as file:
    file.write(data)

Run it from Terminal, Command Prompt, or PowerShell with Python 3:

python3 extract_webarchive.py "saved-page.webarchive" "saved-page.html"

The output is normally the archive’s main HTML resource. Check the archive’s WebResourceMIMEType if the output is not readable text; the main resource can instead be a PDF, image, or another file type.

Online conversion and desktop alternatives

Safari is suitable for a one-off conversion on macOS, while the Python method is repeatable and works on other systems with Python 3. For a non-sensitive file, CloudConvert or Zamzar can be considered only if the service’s current format list explicitly offers .webarchive as an input and HTML as an output. Uploading an archive can disclose its page contents, embedded data, URLs, or authentication-related information, and online support for this Apple-specific format is inconsistent.

Quality and compatibility limits

  • Extracted HTML may still reference the original website, so images and styles can disappear offline or after the site changes.
  • JavaScript that depends on Safari APIs, cookies, local storage, authentication, server-side requests, or external services may not work after extraction.
  • A complete offline copy requires extracting the subresources listed in the archive’s WebSubresources array, writing them to an asset directory, and rewriting the HTML references. The simple script extracts only WebMainResource.
  • The main resource can represent the original document source rather than the final DOM produced by JavaScript. Verify the result in a browser and inspect its image, stylesheet, and script links.
  • Saving HTML changes the container format; it does not preserve all Safari archive metadata or guarantee pixel-identical rendering.

WEBARCHIVE vs HTML: format comparison

How the WEBARCHIVE and HTML formats compare on the properties that matter most for this conversion.

Comparison of the WEBARCHIVE and HTML file formats
Property .WEBARCHIVE Safari Web Archive .HTML HyperText Markup Language
Opens in a web browser No Yes, natively
Plain-text readable — Yes
Text formatting Basic Extensive
Images Yes Yes
Interactive elements Limited Yes
Further editing Limited Easy
Typical file size Large Small
Open standard No Yes
Introduced 2003 1993
Developer Apple Inc. W3C / WHATWG
MIME type application/x-webarchive text/html

Additional formats for
webarchive file conversion

Convert to html from
other formats

Share on social media: