AI in Practice Survey 2025

AI is being built everywhere—but the real question is where it's working in production. We wanted to move past demos and understand how teams are actually deploying durable, enterprise-grade AI systems. "Vibe-coding" is great for exploration, but production AI requires a very different operating model.

We surveyed 413 technical builders to map where adoption is happening, where the gaps are, and how teams are hardening and scaling AI in practice. Here's what we found.

Key Findings

1

Builders are Doing All the Things

Respondents are in the spaghetti phase - they are trying everything. We expected a greater split of techniques and approaches, and instead what we're seeing is that folks that are building in AI are trying it all, with the lowest adoption so far with MCPs.

View supporting data →
2

Open Source is Dominating

We expected fully closed source shops to make up far more of the population - we were wrong. A hybrid reality is winning: openness for control/cost; selective closed-source where latency, capability, or compliance demands it.

View supporting data →
3

AI is Fueling the Internal Performance Engine

As much as we hear about products built on top of AI, we see a huge uptake in AI applications to internal products.

View supporting data →
4

Agents are Going Deeper than API Calling

Deep systems access is table stakes. Agents most commonly use database access, web search, and memory systems.

View supporting data →
5

Evaluations are Prevalent

Over 99% of respondents are assessing quality, and predominantly with AI evaluations. Everyone is doing evaluations now, both manual and automated. We see a broad array of techniques for automation – regression tests, verification, and error analyses. These evaluations are spread surprisingly evenly across dev stage.

View supporting data →
6

RLFT Delivers Significant Lift

Among RLFT users, 80%+ report lifts above 16%, and nearly a third see >30%. That is not incremental improvement — that is category-defining performance delta. If your competitors are doing RLFT and you are not, you are now behind.

View supporting data →
7

Fine-Tuning is Mainstream, and OpenAI is Winning

Overall, 80% of respondents are fine-tuning however this is driven by larger enterprises. 52.4% of start-ups are not (versus 17% for larger companies). This implies that early-stage buyers need out-of-the-box quality and lighter PromptOps; heavy post-training is a mid-to-large enterprise muscle.

View supporting data →
8

MCP is Crossing the Chasm

MCPs have the lowest adoption of the techniques probed in our survey, however a third using LLM chat clients to access data are doing so via MCP. Folks using MCPs are delivering on internal & external projects, technical & non-technical users.

View supporting data →
9

Synthetic Data Powers Evals

63%+ of respondents are using synthetic data for their evaluations. Expect a near-term surge in eval-data marketplaces, scenario libraries, and failure-mode corpora.

View supporting data →

Filter Responses

413 of 413

Survey Results

Age Range

413 responses

Age Range — 413 responses
ResponseDisplayed result
1997-2007,Generation Z15.3%
1981-1996, Millennials61.3%
1965-1980, Generation X22.3%
1946-1964, Baby Boomers1.2%
1997-2007,Generation Z15.3%
1981-1996, Millennials61.3%
1965-1980, Generation X22.3%
1946-1964, Baby Boomers1.2%

Region

413 responses

Region — 413 responses
ResponseDisplayed result
Southeast23.7%
Northeast22.8%
West18.4%
Southwest16.2%
Midwest16%
NA2.9%
Southeast23.7%
Northeast22.8%
West18.4%
Southwest16.2%
Midwest16%
NA2.9%

Company Scale

413 responses

Company Scale — 413 responses
ResponseDisplayed result
1001+27.1%
201 - 100037.3%
11 - 20030.5%
1- 105.1%
1001+27.1%
201 - 100037.3%
11 - 20030.5%
1- 105.1%

Sector

413 responses

Sector — 413 responses
ResponseDisplayed result
Technology & Software49.9%
Manufacturing, Energy & Industrial14.5%
Financial Services & Insurance11.1%
Healthcare & Life Sciences9.7%
Government & Public Sector6.5%
Retail, eCommerce & Consumer Goods4.1%
Media, Telecom & Entertainment2.4%
Other1.5%
Medical Devices0.2%
Technology & Software49.9%
Manufacturing, Energy & Industrial14.5%
Financial Services & Insurance11.1%
Healthcare & Life Sciences9.7%
Government & Public Sector6.5%
Retail, eCommerce & Consumer Goods4.1%
Media, Telecom & Entertainment2.4%
Other1.5%
Medical Devices0.2%

What kinds of AI products are you building?

413 responses

What kinds of AI products are you building? — 413 responses
ResponseDisplayed result
Internal AI tools/workflows65.6%
AI for improving existing products61.3%
Internal user-facing AI features57.1%
Systems for enabling AI products54%
External user-facing AI features51.3%
MCP servers16.7%
Internal AI tools/workflows65.6%
AI for improving existing products61.3%
Internal user-facing AI features57.1%
Systems for enabling AI products54%
External user-facing AI features51.3%
MCP servers16.7%

When in the product lifecycle are you thinking about evals?

413 responses

When in the product lifecycle are you thinking about evals? — 413 responses
ResponseDisplayed result
After deployment, based on user feedback (automated or unstructured)27.1%
During the product spec/definition26.2%
During the QA phase before deployment25.2%
During private alphas and design partnerships19.1%
N/A: We are not considering evals at this stage2.4%
After deployment, based on user feedback (automated or unstructured)27.1%
During the product spec/definition26.2%
During the QA phase before deployment25.2%
During private alphas and design partnerships19.1%
N/A: We are not considering evals at this stage2.4%

What kind of tools are your agents using?

413 responses

What kind of tools are your agents using? — 413 responses
ResponseDisplayed result
Database access72.4%
Web search59.1%
Memory systems55.2%
File systems55.2%
Code interpreter45.8%
First party API/MCP servers45.5%
Plotting/dashboarding tools37.8%
Other (please specify)0.2%
Database access72.4%
Web search59.1%
Memory systems55.2%
File systems55.2%
Code interpreter45.8%
First party API/MCP servers45.5%
Plotting/dashboarding tools37.8%
Other (please specify)0.2%

How are you storing interactions with your AI system?

413 responses

How are you storing interactions with your AI system? — 413 responses
ResponseDisplayed result
We are storing everything as traces in a product designed for LLM observability56.4%
We are using a database and logging traces as rows52.1%
We are storing user prompts and feedback but not the full traces40.2%
We are storing the responses of tool calls39.7%
N/A: We are not storing these data3.1%
Other (please specify)0.2%
We are storing everything as traces in a product designed for LLM observability56.4%
We are using a database and logging traces as rows52.1%
We are storing user prompts and feedback but not the full traces40.2%
We are storing the responses of tool calls39.7%
N/A: We are not storing these data3.1%
Other (please specify)0.2%

How are you reviewing the data of interactions with your system?

413 responses

How are you reviewing the data of interactions with your system? — 413 responses
ResponseDisplayed result
We use spreadsheets to look at the data and extract insights57.1%
We use a specialized LLM observability tool53%
We use a traditional Business Intelligence/Data Science tool51.3%
N/A: We do not currently review the data as part of our process1.9%
We use spreadsheets to look at the data and extract insights57.1%
We use a specialized LLM observability tool53%
We use a traditional Business Intelligence/Data Science tool51.3%
N/A: We do not currently review the data as part of our process1.9%

How are you assessing the quality of your AI products?

413 responses

How are you assessing the quality of your AI products? — 413 responses
ResponseDisplayed result
AI evaluations (manual or automated)66.6%
User satisfaction surveys60%
A/B testing with users47.7%
Subjective qualitative judgment (i.e., "Vibes")41.2%
Telemetry24%
Other (please specify)1.7%
N/A: We aren't yet looking at quality1.2%
AI evaluations (manual or automated)66.6%
User satisfaction surveys60%
A/B testing with users47.7%
Subjective qualitative judgment (i.e., "Vibes")41.2%
Telemetry24%
Other (please specify)1.7%
N/A: We aren't yet looking at quality1.2%

If you're doing AI evaluations, are you...

413 responses

If you're doing AI evaluations, are you... — 413 responses
ResponseDisplayed result
Using human feedback73.6%
Using automated evals63.4%
Using LLM-as-a-judge40.4%
Other (please specify)1.7%
N/A1.2%
Using human feedback73.6%
Using automated evals63.4%
Using LLM-as-a-judge40.4%
Other (please specify)1.7%
N/A1.2%

If you're doing human-feedback, are you...

413 responses

If you're doing human-feedback, are you... — 413 responses
ResponseDisplayed result
A combination of internal and external reviewers51.3%
Have internal experts reviewing the results29.1%
Working with data labeling companies17.9%
N/A: We are not doing human-feedback1.7%
A combination of internal and external reviewers51.3%
Have internal experts reviewing the results29.1%
Working with data labeling companies17.9%
N/A: We are not doing human-feedback1.7%

If you're using automated evals, are you...

413 responses

If you're using automated evals, are you... — 413 responses
ResponseDisplayed result
Using verification strategies52.5%
Doing error analysis on the failure cases50.8%
Running them as regression tests49.4%
Running them as unit tests40.2%
N/A8.5%
Using verification strategies52.5%
Doing error analysis on the failure cases50.8%
Running them as regression tests49.4%
Running them as unit tests40.2%
N/A8.5%

Are you using open source or closed source models for AI capabilities your team is building?

413 responses

Are you using open source or closed source models for AI capabilities your team is building? — 413 responses
ResponseDisplayed result
Only open source models21.8%
Mostly open source, but some closed source44.6%
Even split8.2%
Mostly closed source, but some open source10.2%
Only closed source models15.3%
Only open source models21.8%
Mostly open source, but some closed source44.6%
Even split8.2%
Mostly closed source, but some open source10.2%
Only closed source models15.3%

Have you tried fine tuning?

413 responses

Have you tried fine tuning? — 413 responses
ResponseDisplayed result
Yes, via OpenAI sft product54.5%
Yes, via another paid service19.9%
Yes, via a homegrown library6.8%
No18.9%
Yes, via OpenAI sft product54.5%
Yes, via another paid service19.9%
Yes, via a homegrown library6.8%
No18.9%

Does your AI application/agent have a reward function that you are optimizing for?

413 responses

Does your AI application/agent have a reward function that you are optimizing for? — 413 responses
ResponseDisplayed result
Yes, we specify this reward function mathematically33.4%
Yes, we specify this reward function via natural language and use an LLM judge36.3%
Yes, we evaluate this reward function via human ratings13.8%
No, we don’t improve our agent via a clear reward function16.5%
Yes, we specify this reward function mathematically33.4%
Yes, we specify this reward function via natural language and use an LLM judge36.3%
Yes, we evaluate this reward function via human ratings13.8%
No, we don’t improve our agent via a clear reward function16.5%

If you're using RLFT, how much lift on your evaluation function are you seeing?

413 responses

If you're using RLFT, how much lift on your evaluation function are you seeing? — 413 responses
ResponseDisplayed result
1-15%7.5%
16-30%45.8%
31-45%27.1%
>45%3.4%
We are not using RLFT16.2%
1-15%7.5%
16-30%45.8%
31-45%27.1%
>45%3.4%
We are not using RLFT16.2%

Where are you currently iterating on prompts?

412 responses

Where are you currently iterating on prompts? — 412 responses
ResponseDisplayed result
Prompts are managed as code; engineers check them into our code repository after a normal code review process52.7%
Prompts are managed outside the code base; they are text files that people from different teams collaborate on40.8%
Prompts aren’t designed by humans; they’re purely automated after the first version6.6%
Prompts are managed as code; engineers check them into our code repository after a normal code review process52.7%
Prompts are managed outside the code base; they are text files that people from different teams collaborate on40.8%
Prompts aren’t designed by humans; they’re purely automated after the first version6.6%

Are you using any automated methods for improving context?

413 responses

Are you using any automated methods for improving context? — 413 responses
ResponseDisplayed result
We are using a prompt optimization method (GEPA, PromptEvolution, etc.)47.2%
We are doing ablations on prompt components with automated evals33.7%
We are fully manual in prompt construction19.1%
We are using a prompt optimization method (GEPA, PromptEvolution, etc.)47.2%
We are doing ablations on prompt components with automated evals33.7%
We are fully manual in prompt construction19.1%

Are you using synthetic data generation?

413 responses

Are you using synthetic data generation? — 413 responses
ResponseDisplayed result
Yes, we use it for generating evals63%
Yes, we use it for fine-tuning/post-training22.3%
No14.8%
Yes, we use it for generating evals63%
Yes, we use it for fine-tuning/post-training22.3%
No14.8%

Are you using LLM chat clients to access APIs or data sources via MCP or direct tool calling?

413 responses

Are you using LLM chat clients to access APIs or data sources via MCP or direct tool calling? — 413 responses
ResponseDisplayed result
API via tools36.3%
MCP32.9%
A2A17.2%
No current usage13.6%
API via tools36.3%
MCP32.9%
A2A17.2%
No current usage13.6%

If you're using MCPs, are you...

136 responses (from MCP users)

If you're using MCPs, are you... — 136 responses (from MCP users)
ResponseDisplayed result
Building internal MCPs to access external APIs61%
Using first party MCPs published by other software providers51.5%
Building internal MCPs to access data in your own systems50.7%
Building internal MCPs to access external APIs61%
Using first party MCPs published by other software providers51.5%
Building internal MCPs to access data in your own systems50.7%

If you're building your own MCPs, are you...

136 responses (from MCP users)

If you're building your own MCPs, are you... — 136 responses (from MCP users)
ResponseDisplayed result
Building them to connect to your data warehouse55.1%
Building them to connect to your customer data50%
Building them to connect to your internal tools45.6%
Building them to connect to your data warehouse55.1%
Building them to connect to your customer data50%
Building them to connect to your internal tools45.6%

If you're building MCPs, who is your intended audience?

136 responses (from MCP users)

If you're building MCPs, who is your intended audience? — 136 responses (from MCP users)
ResponseDisplayed result
Internal technical users (e.g. engineers)57.4%
Internal non-technical users (e.g. business analysts)53.7%
External users (e.g. consumers)52.9%
Internal technical users (e.g. engineers)57.4%
Internal non-technical users (e.g. business analysts)53.7%
External users (e.g. consumers)52.9%

Where are you experiencing the biggest pain points today in deploying your AI products? (Ranked 1-7, 1=least painful, 7=most painful)

Average ranking (1=least painful, 7=most painful)

Where are you experiencing the biggest pain points today in deploying your AI products? (Ranked 1-7, 1=least painful, 7=most painful) — Average ranking (1=least painful, 7=most painful)
ResponseDisplayed result
Data quality assessment (Evals)Avg: 3.67
Security & governanceAvg: 3.81
Data review of interactions with our production AIAvg: 3.83
Model managementAvg: 4.09
Storage of interactions with our AIAvg: 4.10
User acceptanceAvg: 4.10
ContextAvg: 4.59
Data quality assessment (Evals)Avg: 3.67
Security & governanceAvg: 3.81
Data review of interactions with our production AIAvg: 3.83
Model managementAvg: 4.09
Storage of interactions with our AIAvg: 4.10
User acceptanceAvg: 4.10
ContextAvg: 4.59

Which area of tooling is delivering the most value today for your projects?

411 responses

Which area of tooling is delivering the most value today for your projects? — 411 responses
ResponseDisplayed result
Data review of interactions with our production AI20.9%
Data quality assessment (evals)19.2%
Security & governance12.9%
Storage of interactions with our AI12.7%
Model management12.4%
User acceptance11.7%
Context7.8%
Data Quality Assessment (Evals)2.4%
Data review of interactions with our production AI20.9%
Data quality assessment (evals)19.2%
Security & governance12.9%
Storage of interactions with our AI12.7%
Model management12.4%
User acceptance11.7%
Context7.8%
Data Quality Assessment (Evals)2.4%

Which area of tooling are you most excited about for the value it may create in the future for your projects?

413 responses

Which area of tooling are you most excited about for the value it may create in the future for your projects? — 413 responses
ResponseDisplayed result
Data review of interactions with our production AI21.3%
Data quality assessment (evals)21.1%
Storage of interactions with our AI17.9%
Security & governance10.2%
User acceptance10.2%
Model management8.5%
Context7.3%
Data Quality Assessment (Evals)3.6%
Data review of interactions with our production AI21.3%
Data quality assessment (evals)21.1%
Storage of interactions with our AI17.9%
Security & governance10.2%
User acceptance10.2%
Model management8.5%
Context7.3%
Data Quality Assessment (Evals)3.6%