AI in Practice Survey 2025
AI is being built everywhere—but the real question is where it's working in production. We wanted to move past demos and understand how teams are actually deploying durable, enterprise-grade AI systems. "Vibe-coding" is great for exploration, but production AI requires a very different operating model.
We surveyed 413 technical builders to map where adoption is happening, where the gaps are, and how teams are hardening and scaling AI in practice. Here's what we found.
Key Findings
Builders are Doing All the Things
Respondents are in the spaghetti phase - they are trying everything. We expected a greater split of techniques and approaches, and instead what we're seeing is that folks that are building in AI are trying it all, with the lowest adoption so far with MCPs.
Open Source is Dominating
We expected fully closed source shops to make up far more of the population - we were wrong. A hybrid reality is winning: openness for control/cost; selective closed-source where latency, capability, or compliance demands it.
AI is Fueling the Internal Performance Engine
As much as we hear about products built on top of AI, we see a huge uptake in AI applications to internal products.
Agents are Going Deeper than API Calling
Deep systems access is table stakes. Agents most commonly use database access, web search, and memory systems.
Evaluations are Prevalent
Over 99% of respondents are assessing quality, and predominantly with AI evaluations. Everyone is doing evaluations now, both manual and automated. We see a broad array of techniques for automation – regression tests, verification, and error analyses. These evaluations are spread surprisingly evenly across dev stage.
RLFT Delivers Significant Lift
Among RLFT users, 80%+ report lifts above 16%, and nearly a third see >30%. That is not incremental improvement — that is category-defining performance delta. If your competitors are doing RLFT and you are not, you are now behind.
Not fine-tuning
Fine-tuning
Not
Fine-tuning
Fine-Tuning is Mainstream, and OpenAI is Winning
Overall, 80% of respondents are fine-tuning however this is driven by larger enterprises. 52.4% of start-ups are not (versus 17% for larger companies). This implies that early-stage buyers need out-of-the-box quality and lighter PromptOps; heavy post-training is a mid-to-large enterprise muscle.
MCP is Crossing the Chasm
MCPs have the lowest adoption of the techniques probed in our survey, however a third using LLM chat clients to access data are doing so via MCP. Folks using MCPs are delivering on internal & external projects, technical & non-technical users.
Synthetic Data Powers Evals
63%+ of respondents are using synthetic data for their evaluations. Expect a near-term surge in eval-data marketplaces, scenario libraries, and failure-mode corpora.
Filter Responses
Survey Results
Age Range
413 responses
| Response | Displayed result |
|---|---|
| 1997-2007,Generation Z | 15.3% |
| 1981-1996, Millennials | 61.3% |
| 1965-1980, Generation X | 22.3% |
| 1946-1964, Baby Boomers | 1.2% |
Region
413 responses
| Response | Displayed result |
|---|---|
| Southeast | 23.7% |
| Northeast | 22.8% |
| West | 18.4% |
| Southwest | 16.2% |
| Midwest | 16% |
| NA | 2.9% |
Company Scale
413 responses
| Response | Displayed result |
|---|---|
| 1001+ | 27.1% |
| 201 - 1000 | 37.3% |
| 11 - 200 | 30.5% |
| 1- 10 | 5.1% |
Sector
413 responses
| Response | Displayed result |
|---|---|
| Technology & Software | 49.9% |
| Manufacturing, Energy & Industrial | 14.5% |
| Financial Services & Insurance | 11.1% |
| Healthcare & Life Sciences | 9.7% |
| Government & Public Sector | 6.5% |
| Retail, eCommerce & Consumer Goods | 4.1% |
| Media, Telecom & Entertainment | 2.4% |
| Other | 1.5% |
| Medical Devices | 0.2% |
What kinds of AI products are you building?
413 responses
| Response | Displayed result |
|---|---|
| Internal AI tools/workflows | 65.6% |
| AI for improving existing products | 61.3% |
| Internal user-facing AI features | 57.1% |
| Systems for enabling AI products | 54% |
| External user-facing AI features | 51.3% |
| MCP servers | 16.7% |
When in the product lifecycle are you thinking about evals?
413 responses
| Response | Displayed result |
|---|---|
| After deployment, based on user feedback (automated or unstructured) | 27.1% |
| During the product spec/definition | 26.2% |
| During the QA phase before deployment | 25.2% |
| During private alphas and design partnerships | 19.1% |
| N/A: We are not considering evals at this stage | 2.4% |
What kind of tools are your agents using?
413 responses
| Response | Displayed result |
|---|---|
| Database access | 72.4% |
| Web search | 59.1% |
| Memory systems | 55.2% |
| File systems | 55.2% |
| Code interpreter | 45.8% |
| First party API/MCP servers | 45.5% |
| Plotting/dashboarding tools | 37.8% |
| Other (please specify) | 0.2% |
How are you storing interactions with your AI system?
413 responses
| Response | Displayed result |
|---|---|
| We are storing everything as traces in a product designed for LLM observability | 56.4% |
| We are using a database and logging traces as rows | 52.1% |
| We are storing user prompts and feedback but not the full traces | 40.2% |
| We are storing the responses of tool calls | 39.7% |
| N/A: We are not storing these data | 3.1% |
| Other (please specify) | 0.2% |
How are you reviewing the data of interactions with your system?
413 responses
| Response | Displayed result |
|---|---|
| We use spreadsheets to look at the data and extract insights | 57.1% |
| We use a specialized LLM observability tool | 53% |
| We use a traditional Business Intelligence/Data Science tool | 51.3% |
| N/A: We do not currently review the data as part of our process | 1.9% |
How are you assessing the quality of your AI products?
413 responses
| Response | Displayed result |
|---|---|
| AI evaluations (manual or automated) | 66.6% |
| User satisfaction surveys | 60% |
| A/B testing with users | 47.7% |
| Subjective qualitative judgment (i.e., "Vibes") | 41.2% |
| Telemetry | 24% |
| Other (please specify) | 1.7% |
| N/A: We aren't yet looking at quality | 1.2% |
If you're doing AI evaluations, are you...
413 responses
| Response | Displayed result |
|---|---|
| Using human feedback | 73.6% |
| Using automated evals | 63.4% |
| Using LLM-as-a-judge | 40.4% |
| Other (please specify) | 1.7% |
| N/A | 1.2% |
If you're doing human-feedback, are you...
413 responses
| Response | Displayed result |
|---|---|
| A combination of internal and external reviewers | 51.3% |
| Have internal experts reviewing the results | 29.1% |
| Working with data labeling companies | 17.9% |
| N/A: We are not doing human-feedback | 1.7% |
If you're using automated evals, are you...
413 responses
| Response | Displayed result |
|---|---|
| Using verification strategies | 52.5% |
| Doing error analysis on the failure cases | 50.8% |
| Running them as regression tests | 49.4% |
| Running them as unit tests | 40.2% |
| N/A | 8.5% |
Are you using open source or closed source models for AI capabilities your team is building?
413 responses
| Response | Displayed result |
|---|---|
| Only open source models | 21.8% |
| Mostly open source, but some closed source | 44.6% |
| Even split | 8.2% |
| Mostly closed source, but some open source | 10.2% |
| Only closed source models | 15.3% |
Have you tried fine tuning?
413 responses
| Response | Displayed result |
|---|---|
| Yes, via OpenAI sft product | 54.5% |
| Yes, via another paid service | 19.9% |
| Yes, via a homegrown library | 6.8% |
| No | 18.9% |
Does your AI application/agent have a reward function that you are optimizing for?
413 responses
| Response | Displayed result |
|---|---|
| Yes, we specify this reward function mathematically | 33.4% |
| Yes, we specify this reward function via natural language and use an LLM judge | 36.3% |
| Yes, we evaluate this reward function via human ratings | 13.8% |
| No, we don’t improve our agent via a clear reward function | 16.5% |
If you're using RLFT, how much lift on your evaluation function are you seeing?
413 responses
| Response | Displayed result |
|---|---|
| 1-15% | 7.5% |
| 16-30% | 45.8% |
| 31-45% | 27.1% |
| >45% | 3.4% |
| We are not using RLFT | 16.2% |
Where are you currently iterating on prompts?
412 responses
| Response | Displayed result |
|---|---|
| Prompts are managed as code; engineers check them into our code repository after a normal code review process | 52.7% |
| Prompts are managed outside the code base; they are text files that people from different teams collaborate on | 40.8% |
| Prompts aren’t designed by humans; they’re purely automated after the first version | 6.6% |
Are you using any automated methods for improving context?
413 responses
| Response | Displayed result |
|---|---|
| We are using a prompt optimization method (GEPA, PromptEvolution, etc.) | 47.2% |
| We are doing ablations on prompt components with automated evals | 33.7% |
| We are fully manual in prompt construction | 19.1% |
Are you using synthetic data generation?
413 responses
| Response | Displayed result |
|---|---|
| Yes, we use it for generating evals | 63% |
| Yes, we use it for fine-tuning/post-training | 22.3% |
| No | 14.8% |
Are you using LLM chat clients to access APIs or data sources via MCP or direct tool calling?
413 responses
| Response | Displayed result |
|---|---|
| API via tools | 36.3% |
| MCP | 32.9% |
| A2A | 17.2% |
| No current usage | 13.6% |
If you're using MCPs, are you...
136 responses (from MCP users)
| Response | Displayed result |
|---|---|
| Building internal MCPs to access external APIs | 61% |
| Using first party MCPs published by other software providers | 51.5% |
| Building internal MCPs to access data in your own systems | 50.7% |
If you're building your own MCPs, are you...
136 responses (from MCP users)
| Response | Displayed result |
|---|---|
| Building them to connect to your data warehouse | 55.1% |
| Building them to connect to your customer data | 50% |
| Building them to connect to your internal tools | 45.6% |
If you're building MCPs, who is your intended audience?
136 responses (from MCP users)
| Response | Displayed result |
|---|---|
| Internal technical users (e.g. engineers) | 57.4% |
| Internal non-technical users (e.g. business analysts) | 53.7% |
| External users (e.g. consumers) | 52.9% |
Where are you experiencing the biggest pain points today in deploying your AI products? (Ranked 1-7, 1=least painful, 7=most painful)
Average ranking (1=least painful, 7=most painful)
| Response | Displayed result |
|---|---|
| Data quality assessment (Evals) | Avg: 3.67 |
| Security & governance | Avg: 3.81 |
| Data review of interactions with our production AI | Avg: 3.83 |
| Model management | Avg: 4.09 |
| Storage of interactions with our AI | Avg: 4.10 |
| User acceptance | Avg: 4.10 |
| Context | Avg: 4.59 |
Which area of tooling is delivering the most value today for your projects?
411 responses
| Response | Displayed result |
|---|---|
| Data review of interactions with our production AI | 20.9% |
| Data quality assessment (evals) | 19.2% |
| Security & governance | 12.9% |
| Storage of interactions with our AI | 12.7% |
| Model management | 12.4% |
| User acceptance | 11.7% |
| Context | 7.8% |
| Data Quality Assessment (Evals) | 2.4% |
Which area of tooling are you most excited about for the value it may create in the future for your projects?
413 responses
| Response | Displayed result |
|---|---|
| Data review of interactions with our production AI | 21.3% |
| Data quality assessment (evals) | 21.1% |
| Storage of interactions with our AI | 17.9% |
| Security & governance | 10.2% |
| User acceptance | 10.2% |
| Model management | 8.5% |
| Context | 7.3% |
| Data Quality Assessment (Evals) | 3.6% |
- n=413 practitioners at companies with meaningful AI build efforts
- Weighted to mid-market and enterprise companies, with only 5% of respondents at early startups.
- Most respondents come from the technology & software sector, with representation from manufacturing, fintech, healthcare, & government.
- Respondents are predominantly technical leaders/builders (directly building or managing teams that build AI).