Skip to main content
New tool CRON Expression Builder — preview next run times before you schedule Apex. Open the builder →
A 3D glass cube representing data nodes for optimizing stateful batch apex performance
Apex

Stateful Batch Apex: Should You Cache Picklist Values?

Holding metadata like picklist values in stateful variables looks like a free performance win, and it brings serialization limits, heap pressure and stale data along with it. Here is when the trade is worth making and when it is not.

Key takeaways Reserve Database.Stateful for variables that track the state of the batch execution itself, such as row counts or processed IDs. Schema.DescribeFieldResult is highly efficient. Do not treat it like a traditional SOQL query that needs manual caching. Storing large collections in stateful variables adds serialization and deserialization time to every single batch chunk, which can push total execution time up rather than down. For per-chunk optimization, a static variable in a helper class stops the re-fetch without persisting state across transactions.

The allure of the stateful variable

Every Database.Batchable class forces a choice between stateless and stateful execution. The Database.Stateful interface keeps the values of instance variables alive across execution chunks, and for a lot of developers that is a siren song. Why query the same metadata, picklist values or custom settings, a thousand times when you can query it once in the start method and hold it in a variable for the duration of the job?

The performance benefit of minimizing SOQL queries is clear. So are the risks to data integrity and heap size limits, and that is the trade you are actually making.

The mechanics of stateful batch

When a batch class implements Database.Stateful, Salesforce serializes the object's instance variables at the end of every execute method and deserializes them at the beginning of the next.

Cache picklist values, perhaps by pulling them from Schema.DescribeFieldResult into a Map<String, List<String>>, and you have made the serialized footprint of that object bigger. Here is the standard way developers typically attempt this:

global class PicklistCacheBatch implements Database.Batchable<SObject>, Database.Stateful {
    private Map<String, List<String>> picklistCache = new Map<String, List<String>>();

    global PicklistCacheBatch() {
        // Populating the cache once
        Schema.DescribeFieldResult fieldResult = Account.Industry.getDescribe();
        List<Schema.PicklistEntry> ple = fieldResult.getPicklistValues();
        List<String> values = new List<String>();
        for(Schema.PicklistEntry f : ple) {
            values.add(f.getValue());
        }
        picklistCache.put('Industry', values);
    }

    global void execute(Database.BatchableContext BC, List<Account> scope) {
        // Using the stateful map
        for(Account a : scope) {
            if(picklistCache.get('Industry').contains(a.Industry)) {
                // Process logic
            }
        }
    }
}

It looks efficient. Now think about what happens when picklistCache grows large. Every time a batch chunk finishes, that entire map is serialized into the database. On large batches with high concurrency you can hit heap size or serialization limits that cripple the job before it finishes.

The hidden costs: memory and serialization

Three things make caching metadata in a Stateful variable an anti-pattern most of the time:

  1. Salesforce has strict limits on the size of the stateful object. If your picklist data is massive, or you are caching many fields, the object exceeds the serialization limit and you get a "Serializing stateful batch failed" error.
  2. The object persists across execution chunks, so its memory stays allocated for the lifetime of the batch. That eats into your total heap allocation and can leave less room for the actual processing logic inside execute.
  3. Batch jobs can take a long time to complete. If a metadata change lands while the batch is running, say someone deactivates a picklist value, your cached map now holds stale, incorrect information.

When is caching actually appropriate?

If standard schema calls really are too heavy for your job, look at the alternatives first. Schema methods are generally very fast, because the platform caches them at the system level.

In most scenarios, calling getDescribe() inside the execute method is safer and cheaper than many developers realize. The platform overhead of that call is often lower than the overhead of serializing and deserializing a large stateful object.

If you are performing complex calculations on thousands of picklist values per record, though, a static map is worth considering ahead of a stateful instance variable:

public class PicklistHelper {
    private static Map<String, List<String>> cachedMap;

    public static List<String> getIndustryValues() {
        if (cachedMap == null) {
            cachedMap = new Map<String, List<String>>();
            // Fill the map
        }
        return cachedMap.get('Industry');
    }
}

A static variable in a separate utility class gives you a lazy load. It caches the data for the duration of the transaction, the current chunk, which is often exactly what you need without the baggage of Database.Stateful.

Recommendations

For the vast majority of Salesforce implementations, keep picklist values and other metadata out of Database.Stateful variables. The platform's built-in schema caching is already doing this work for you.

  • Use Database.Stateful for counters, error logs, or aggregate summaries that strictly have to pass between chunks.
  • If performance is a concern, cache in a static utility class so the data lives for the execution scope of a single chunk.
  • If you are worried about the cost of your Describe calls, measure it. System.Limits.getHeapSize() in your test classes will show the impact.
Newsletter

One email every Tuesday

New guides, tool updates, and the release-note changes that break things.

No spam. Unsubscribe in one click.

Comments

Loading comments...

Leave a Comment