Skip to main content
New tool CRON Expression Builder — preview next run times before you schedule Apex. Open the builder →
3D abstract of complex data pipeline transforming into a streamlined, glowing circuit, representing reduced development time.
Agentforce & AI

AI Data Integration: Informatica's Pipeline Dev Time Reduction

Informatica used AI and prompt engineering to cut data integration pipeline development from days to minutes. Their team shares the technical problems they hit building trustworthy AI-assisted integration.

Key takeaways Informatica cut data integration pipeline development time from days to minutes by adopting AI and natural language interfaces. Traditional data integration was bottlenecked by complexity; working from user intent removes most of it. Reliable AI integration meant adapting to fast-moving LLMs and re-evaluating how development and testing work. Accuracy in AI-generated pipelines comes from prompt engineering, contextual grounding and thorough validation. What's next is broader transformation support, code-first experiences and agentic workflows, without dropping the accuracy and reliability bar.

Informatica's AI-Driven Approach to Data Integration Pipeline Development

Informatica moved data integration pipeline development from a multi-day job to a few minutes, mostly by putting AI and prompt engineering in front of it. Their engineering team wrote up what it took to build Copilot, the AI-powered experience that generates pipelines from natural language, and the problems they hit are worth reading if you work anywhere near this.

The bottleneck of traditional data integration

The old workflow made engineers do the boring parts by hand. You inspected metadata, configured transformations, and wired sources to targets one field at a time. Even a simple pipeline could eat days or weeks, because the complexity was real and you had to know a long list of transformation types and configuration options to get through it. Informatica's read is that the integration capabilities were always there; what hurt was the user experience wrapped around them.

That pushed them to rebuild the interaction model around natural language. Rather than walking a user through a pipeline step by step, let them describe what they want conversationally and generate the pipeline in minutes. The framing puts the business outcome first and the implementation details second, so users spend their time on the data instead of on the machinery that moves it.

Challenges in transformation design

Hand-building mappings between systems like Snowflake and Salesforce is slow work. You inspect the schemas to learn the structure and relationships on both sides, map fields explicitly, configure the transformation logic, then validate and troubleshoot the output, repeatedly. A small change could set off another round of validation and debugging. Describing the intent conversationally, which is what Copilot does, hides most of that and moves the conversation to the result you actually want.

Customer productivity gains

Early adoption patterns and feature usage gave Informatica their evidence:

  • Customers have generated roughly 10,000 pipelines from conversational prompts since launch.
  • Expression generation, where you describe the behavior instead of writing syntax, drew over 25,000 requests, and about 60% were accepted with no edits.
  • Mapping augmentation, which inserts transformations into an existing pipeline, was used on over 900 pipelines within two months of release, with roughly 80% of generated transformations accepted.

The acceptance rates are the interesting part. Anything new looks fine on a launch chart; thousands of customers coming back to it is a different signal.

Engineering challenges in reliable AI-assisted data integration

Informatica's main problem was that AI moved faster than their architecture could adapt. They started by fine-tuning their own models to meet the project's requirements. As foundation models matured, OpenAI's in particular, those began delivering better outcomes, so they switched. That bought accuracy and cost them control, since more of the behavior now lived in someone else's model.

It also forced a rethink of how they develop and test. Deterministic software gives you the same output for the same input; probabilistic AI does not. A small variation in a prompt can produce a very different response, so test coverage has to get wider to catch regressions. What they landed on was continuous reevaluation, heavy investment in testing, and a focus on prompt engineering and grounding that plays to what large language models are good at.

Accuracy with conversational prompts

Generating a pipeline was the easy half. Making it correct was the harder engineering problem:

  • Metadata discovery and schema mapping. Users often don't know where the data they want lives, or how the schemas line up.
  • Prompt engineering and intent understanding. Turning what someone asked for into pipeline logic that does it.
  • Data quality and output validation. Handling variation in data quality, formatting, missing values and cleansing requirements.

Moving to external LLMs like OpenAI's added its own prompt tuning and grounding problems. Getting enough context in front of the model to steer generation accurately became the job, which meant iterating on prompts and padding them with context. Generated output needed rigorous checking too; expressions and object names occasionally came back needing a fix.

Informatica's answer was to expand cataloging so metadata could be identified automatically, keep refining prompts, ground the LLMs with more contextual data, and add validation layers so what came out was usable. Their conclusion: accuracy improvements came from context, validation and guardrails more than from model size.

Future engineering challenges

Customer expectations keep rising as AI-assisted integration matures. Informatica names three things coming at them: demand for a wider range of transformations and more complex use cases, requests for code-first workflows that behave like an AI coding assistant, and agentic workflows with headless, automated end-to-end development.

The hard part is adding all of that without losing the accuracy, reliability and production readiness people already expect. Trust is what the whole thing rests on, and every new feature has to hold the same bar as the rest of a production system.

Originally reported by engineering.salesforce.com

Newsletter

One email every Tuesday

New guides, tool updates, and the release-note changes that break things.

No spam. Unsubscribe in one click.

Comments

Loading comments...

Leave a Comment