This week Baseten announced that it is one of the first open-model providers in OpenAI’s enterprise marketplace. OpenAI enterprise customers can now run open models directly in Codex or through the Responses API.
It looks like a partnership announcement. It is really a signal that enterprise AI is becoming multi-model, and that even the leading model vendor now expects customers to mix models.
What it changes
The question for leadership is no longer which model to standardize on. It is which model fits each task, at what cost. A planning step may need a frontier model, but a high-volume extraction job usually doesn’t. The teams that match each task to the right model will spend far less than the teams that can’t.
There is a catch. Having access to more models doesn’t make your agents any easier to move. If an agent’s instructions, tools and history are tied to one vendor’s product, switching models means rebuilding the agent. You can only swap models freely if the agent was built so that it doesn’t depend on any one of them.
How you win
Own the agent, not the model. Keep an agent’s instructions, tools and configuration in your own repository, versioned like any other software. The model and the harness should be settings on that agent. Teams that do this can adopt a better or cheaper model the week it ships, while everyone else is still rebuilding.
Test on your own work, not on benchmarks. Public leaderboards tell you very little about how a model handles your codebase, your documents or your customers. Run the same agent on a frontier model and an open model, using a week of real tasks, and compare the cost, latency and quality of each run.
Test the harness too. Codex, Claude Code and a custom LangChain harness can produce very different results with the same model. The harness that looks best in a demo doesn’t always win on your workload.
Put a budget on every experiment. Trying new models and harnesses should be cheap and routine. That only works when every agent has a hard spending limit, so an experiment can never turn into a surprise bill.
Teams that follow these four habits turn every new model release into a chance to cut costs, instead of another migration project.
AgentHippo is one tool that gets you these answers quickly. You can run the same agent across models and harnesses, in your own environment and with your own keys, and see the results side by side within days.
Choosing the model is only half the problem. Once agents act on real systems, the harder question is who each agent is acting for and what it is allowed to do. NVIDIA addressed that this week with its announcement on agent identity and permission enforcement. That is the topic of the next post.
Sources: Baseten announcement · Techmeme