Skip to main content
New tool CRON Expression Builder — preview next run times before you schedule Apex. Open the builder →
3D smartphone showing a large file download progress bar, illustrating robust Salesforce large file downloads.
Integration

Salesforce Large File Downloads: Chunked Integration

Offline-first mobile apps need large file downloads that survive a bad network. How to build chunked, resumable downloads on the Shepherd Servlet and the HTTP Range header.

Key takeaways Large file downloads on unreliable networks require chunking and resumability. The Shepherd Servlet or ContentVersion.VersionData endpoints support HTTP Range, which is what a mobile downloader needs. Maintain a FileDownloadContext to track progress (startByte) and essential file metadata. Treat network failures as recoverable and persist progress, so a download picks up where it stopped. Validate downloaded files using checksums (e.g., MD5) after completion to confirm data integrity.

Handling large Salesforce file downloads in integrations

Downloading large files is where Salesforce integrations for offline-first mobile apps fall over. The standard "one big GET" fails on an unstable network, the progress is gone, and the user is left to start again. What holds up instead is a chunked, resumable download.

Salesforce Files, managed via ContentDocument and ContentVersion objects, can store files up to 2 GB. A file that size needs a download strategy built for the network conditions mobile devices actually see, such as hotel Wi-Fi or roaming cellular data.

Two primary download methods

Salesforce gives you two main ways to retrieve file content.

The Connect Files API is a REST API that reaches Files through endpoints like /connect/files/.... It typically involves fetching file metadata first, then making a separate call to retrieve the binary content. That suits backend integrations on stable networks, though its default streaming pattern makes chunking and resumability for mobile apps complex.

The Shepherd Servlet download endpoint, /sfc/servlet.shepherd/version/download/{ContentVersionId} or /sfc/servlet.shepherd/document/download/{ContentDocumentId}, is the more mobile-friendly option. It supports the HTTP Range header, which is what makes a chunked downloader possible at all.

The ContentVersion/{Id}/VersionData blob-retrieve URL works as the byte endpoint too, and it needs the same client-side logic for chunking and resumption.

Implementing a chunked download strategy

A resilient download breaks the large file into smaller, manageable chunks and makes the process resumable.

Step 1: prepare file context

Before initiating a download, gather the metadata you need into a compact file context object. This object should include:

  • contentDocumentId and contentVersionId: Identifiers for the file.
  • totalSize: The total size of the file in bytes.
  • checksum: An MD5 or similar hash for final file validation.
  • fileName/extension: For local file naming.
  • remoteUrl: The URL to download from (e.g., Shepherd Servlet or VersionData endpoint).
  • downloadingURL: The local temporary path for the downloaded file.
  • startByte: The byte offset to resume from (0 for a new download).
  • chunkSize: The size of each chunk (e.g., 2 to 10 MB).

Step 2: download in chunks with HTTP Range

Chunking runs on the HTTP Range header. For each chunk, send a GET request to the remoteUrl with an Authorization header and a Range header specifying the byte range (e.g., bytes=offset-end).

offset = context.startByte // 0 for new download, >0 if resuming
chunkSize = context.chunkSize
totalSize = context.totalSize

open file at context.downloadingURL in append mode
seek(file, offset)

while offset < totalSize:
    end = min(offset + chunkSize - 1, totalSize - 1)
    response = HTTP GET context.remoteUrl with headers:
    - Authorization: Bearer <token>
    - Range: "bytes=offset-end"

    if response is not successful (timeout, 5xx, 429, etc):
        // Network problem: keep the partial file
        // and remember how far we got
        context.startByte = offset
        save context (status = "paused")
        return "resumable error"

    data = response.body
    write data to file
    offset += data.length
    context.startByte = offset

    // Persist progress so we can resume from here next time
    save context (status = "inProgress")
    report progress = offset / totalSize

close file

After each successful chunk, persist the startByte so you can resume from it. The checksum calculation waits until the whole file is down.

Step 3: resume after network failures

Assume network failures will happen. Rather than discarding partial downloads, treat most failures as recoverable: persist the startByte and mark the context as paused or failed_resumable. When resuming, load the FileDownloadContext, and if startByte > 0 and the partial file still exists, open it in append mode and seek to startByte to continue the download.

Progress then survives both network interruptions and the application being backgrounded.

Step 4: validate the file with MD5

Once the download loop completes (startByte equals totalSize), validate the integrity of the downloaded file using its MD5 checksum.

if context.startByte != context.totalSize:
    return "download not complete"

open file at context.downloadingURL for read
init MD5 calculator

while there is data to read:
    chunk = readNextBlock(file)
    if chunk is empty:
        break
    update MD5 with chunk

computed = finalize MD5

if computed == context.checksum:
    mark context as "completed"
    return success(context.downloadingURL)
else:
    delete file at context.downloadingURL
    delete context entry
    return "checksum mismatch"

Matching checksums mean the file is completed and can move to its final location. A mismatch means deleting the corrupted file and its context, so the next attempt starts fresh.

Connect Files vs. Servlet: optimal use cases

The Connect Files API suits backend services and integrations on stable networks, where complex resume behavior is not a primary concern. It gives you a higher-level REST interface.

The Servlet (or the VersionData blob) with HTTP Range is the one for offline-first mobile applications, where resuming reliably on an unreliable network is the whole point. The cost is managing metadata like size and checksum at the client level.

Plenty of comprehensive solutions run both: Connect Files for the simpler server-side operations, and the Shepherd Servlet or VersionData endpoint with a custom FilesDownloader for mobile scenarios.

Originally reported by salesforceben.com

Newsletter

One email every Tuesday

New guides, tool updates, and the release-note changes that break things.

No spam. Unsubscribe in one click.

Comments

Loading comments...

Leave a Comment