Handling large Salesforce file downloads in integrations
Downloading large files is where Salesforce integrations for offline-first mobile apps fall over. The standard "one big GET" fails on an unstable network, the progress is gone, and the user is left to start again. What holds up instead is a chunked, resumable download.
Salesforce Files, managed via ContentDocument and ContentVersion objects, can store files up to 2 GB. A file that size needs a download strategy built for the network conditions mobile devices actually see, such as hotel Wi-Fi or roaming cellular data.
Two primary download methods
Salesforce gives you two main ways to retrieve file content.
The Connect Files API is a REST API that reaches Files through endpoints like /connect/files/.... It typically involves fetching file metadata first, then making a separate call to retrieve the binary content. That suits backend integrations on stable networks, though its default streaming pattern makes chunking and resumability for mobile apps complex.
The Shepherd Servlet download endpoint, /sfc/servlet.shepherd/version/download/{ContentVersionId} or /sfc/servlet.shepherd/document/download/{ContentDocumentId}, is the more mobile-friendly option. It supports the HTTP Range header, which is what makes a chunked downloader possible at all.
The ContentVersion/{Id}/VersionData blob-retrieve URL works as the byte endpoint too, and it needs the same client-side logic for chunking and resumption.
Implementing a chunked download strategy
A resilient download breaks the large file into smaller, manageable chunks and makes the process resumable.
Step 1: prepare file context
Before initiating a download, gather the metadata you need into a compact file context object. This object should include:
contentDocumentIdandcontentVersionId: Identifiers for the file.totalSize: The total size of the file in bytes.checksum: An MD5 or similar hash for final file validation.fileName/extension: For local file naming.remoteUrl: The URL to download from (e.g., Shepherd Servlet orVersionDataendpoint).downloadingURL: The local temporary path for the downloaded file.startByte: The byte offset to resume from (0 for a new download).chunkSize: The size of each chunk (e.g., 2 to 10 MB).
Step 2: download in chunks with HTTP Range
Chunking runs on the HTTP Range header. For each chunk, send a GET request to the remoteUrl with an Authorization header and a Range header specifying the byte range (e.g., bytes=offset-end).
offset = context.startByte // 0 for new download, >0 if resuming
chunkSize = context.chunkSize
totalSize = context.totalSize
open file at context.downloadingURL in append mode
seek(file, offset)
while offset < totalSize:
end = min(offset + chunkSize - 1, totalSize - 1)
response = HTTP GET context.remoteUrl with headers:
- Authorization: Bearer <token>
- Range: "bytes=offset-end"
if response is not successful (timeout, 5xx, 429, etc):
// Network problem: keep the partial file
// and remember how far we got
context.startByte = offset
save context (status = "paused")
return "resumable error"
data = response.body
write data to file
offset += data.length
context.startByte = offset
// Persist progress so we can resume from here next time
save context (status = "inProgress")
report progress = offset / totalSize
close file
After each successful chunk, persist the startByte so you can resume from it. The checksum calculation waits until the whole file is down.
Step 3: resume after network failures
Assume network failures will happen. Rather than discarding partial downloads, treat most failures as recoverable: persist the startByte and mark the context as paused or failed_resumable. When resuming, load the FileDownloadContext, and if startByte > 0 and the partial file still exists, open it in append mode and seek to startByte to continue the download.
Progress then survives both network interruptions and the application being backgrounded.
Step 4: validate the file with MD5
Once the download loop completes (startByte equals totalSize), validate the integrity of the downloaded file using its MD5 checksum.
if context.startByte != context.totalSize:
return "download not complete"
open file at context.downloadingURL for read
init MD5 calculator
while there is data to read:
chunk = readNextBlock(file)
if chunk is empty:
break
update MD5 with chunk
computed = finalize MD5
if computed == context.checksum:
mark context as "completed"
return success(context.downloadingURL)
else:
delete file at context.downloadingURL
delete context entry
return "checksum mismatch"
Matching checksums mean the file is completed and can move to its final location. A mismatch means deleting the corrupted file and its context, so the next attempt starts fresh.
Connect Files vs. Servlet: optimal use cases
The Connect Files API suits backend services and integrations on stable networks, where complex resume behavior is not a primary concern. It gives you a higher-level REST interface.
The Servlet (or the VersionData blob) with HTTP Range is the one for offline-first mobile applications, where resuming reliably on an unreliable network is the whole point. The cost is managing metadata like size and checksum at the client level.
Plenty of comprehensive solutions run both: Connect Files for the simpler server-side operations, and the Shepherd Servlet or VersionData endpoint with a custom FilesDownloader for mobile scenarios.
Leave a Comment