﻿WEBVTT

1
00:00:09.280 --> 00:00:11.820
<v 0>Okay. Hello, everyone.</v>

2
00:00:13.500 --> 00:00:17.000
I'll wait a little bit for people to settle down. I'm Alex Atallah,

3
00:00:17.300 --> 00:00:22.240
CEO and cofounder of OpenRouter-the largest language model marketplace, router,

4
00:00:22.600 --> 00:00:26.620
and source of public data about what people are doing with language models and

5
00:00:26.680 --> 00:00:30.920
why. So I want to start,

6
00:00:31.140 --> 00:00:34.280
I'll talk a little bit about what OpenRouter is first.

7
00:00:35.520 --> 00:00:36.600
Beginning of 2023,

8
00:00:37.320 --> 00:00:40.960
there was only one real model that people were using,

9
00:00:42.100 --> 00:00:45.080
in January and February, and that was ChatGPT.

10
00:00:45.400 --> 00:00:47.860
It was GPT-3.5 and 3.5 Turbo.

11
00:00:48.880 --> 00:00:51.940
And the open source scene hadn't really developed yet.

12
00:00:52.000 --> 00:00:55.820
Llama came out in January. It was not a chat completions model.

13
00:00:55.880 --> 00:00:57.140
It was a completions model.

14
00:00:57.420 --> 00:00:59.460
So you couldn't really interact with it or chat with it.

15
00:01:00.080 --> 00:01:02.600
But as the ecosystem matured,

16
00:01:02.680 --> 00:01:07.520
it became clear that there were going to be tons of these. And moreover,

17
00:01:07.840 --> 00:01:11.900
that people were going to find really powerful use cases for open source models

18
00:01:12.380 --> 00:01:16.140
that they could optimize for their company and

19
00:01:17.680 --> 00:01:20.940
improve not just cost, but also quality.
And so,

20
00:01:21.300 --> 00:01:25.880
we took a bet and decided to build a marketplace that allowed us to see

21
00:01:26.260 --> 00:01:28.100
all models in one place,

22
00:01:28.600 --> 00:01:31.860
showing everyone across all apps and all use cases,

23
00:01:32.260 --> 00:01:34.420
what people were using AI for and why.

24
00:01:35.340 --> 00:01:38.700
And this is from our rankings page, which is public,

25
00:01:38.980 --> 00:01:43.600
and you can visit it and track it and see all AI usage grouped by model,

26
00:01:43.880 --> 00:01:47.640
grouped by app, grouped by use case, by the type of prompt,

27
00:01:48.060 --> 00:01:49.860
grouped by modality images.

28
00:01:50.260 --> 00:01:53.240
We just launched audio rankings today.

29
00:01:53.580 --> 00:01:57.300
So you can see the top audio input models and which ones are best at doing

30
00:01:57.360 --> 00:01:58.193
speech to text.

31
00:01:59.260 --> 00:02:04.060
And then we've launched video and image and audio output models recently as

32
00:02:04.100 --> 00:02:08.760
well. So all of this data is now public and available not just to humans,

33
00:02:08.820 --> 00:02:13.580
but soon also to agents. So you'll be able to, in Claude Code,

34
00:02:13.820 --> 00:02:18.580
just ask what the most popular models are for your use case and ask

35
00:02:18.820 --> 00:02:20.080
about benchmarks as well,

36
00:02:20.540 --> 00:02:25.160
like which ones are ranking highest for math or highest for science or

37
00:02:25.240 --> 00:02:26.200
highest for coding,

38
00:02:27.100 --> 00:02:31.860
and just iteratively explore the public data in a way that no one has ever done

39
00:02:31.880 --> 00:02:35.880
before.
And that's our vision and that's what we're trying to do with our public

40
00:02:35.940 --> 00:02:37.640
data. So in this presentation,

41
00:02:37.980 --> 00:02:41.440
I'm going to be talking about some insights that we've pulled from our public

42
00:02:41.480 --> 00:02:44.060
data that can help you with pricing.

43
00:02:46.100 --> 00:02:48.860
Nearly everyone's pricing is being challenged.

44
00:02:49.600 --> 00:02:53.240
We see thousands of applications and millions of people using inference,

45
00:02:53.640 --> 00:02:57.640
and everyone is scrambling to figure out how to price their inference-driven

46
00:02:57.740 --> 00:03:01.340
features. You see usage-based pricing, you see a subscription model,

47
00:03:01.660 --> 00:03:05.780
you see hybrids of both, where you subscribe and overage costs double.

48
00:03:06.580 --> 00:03:09.480
You see single-time payments.

49
00:03:10.860 --> 00:03:15.460
People are trying everything and struggling because power users often just use

50
00:03:15.520 --> 00:03:19.600
an enormous amount of inference way more than the average user does.

51
00:03:21.240 --> 00:03:22.073
So

52
00:03:23.220 --> 00:03:26.880
it's easy to feel like AI has wrecked all of our pricing models and that we need

53
00:03:26.960 --> 00:03:28.340
something brand new.

54
00:03:30.340 --> 00:03:33.380
It looks like inference is getting too expensive.

55
00:03:34.660 --> 00:03:38.660
If you consider pay as you go, it's really, really hard to forecast this.

56
00:03:39.880 --> 00:03:44.820
If you're like trying to price in a seat-based way, like agents, like I said,

57
00:03:44.880 --> 00:03:49.000
are power users. They just use 1,000 times what a human uses.

58
00:03:49.000 --> 00:03:52.660
And if you try to price based on outcomes,

59
00:03:52.820 --> 00:03:56.660
which is something that a lot of companies are trying now, it's very difficult.

60
00:03:56.720 --> 00:03:58.260
You assume a lot of risk,

61
00:03:58.760 --> 00:04:02.740
and it's hard to know how your margins are going to scale.

62
00:04:04.840 --> 00:04:05.673
So

63
00:04:06.500 --> 00:04:09.900
I want to take a step back and make you challenge an assumption going in here,

64
00:04:10.040 --> 00:04:14.240
which is that you may be thinking of your inference in an

65
00:04:15.280 --> 00:04:16.140
unscalable way.

66
00:04:17.040 --> 00:04:20.020
It might not be your pricing that you need to think about at all.

67
00:04:20.420 --> 00:04:24.240
It's the way that you're using AI models under the hood and how much they're

68
00:04:24.260 --> 00:04:29.060
costing you. And if you can fix that problem, a lot of pricing problems,

69
00:04:29.920 --> 00:04:33.280
a lot of optionality opens up for you. So

70
00:04:34.880 --> 00:04:39.320
here's some data that we're seeing across the market, across all use cases.

71
00:04:40.280 --> 00:04:43.580
The root problem I'm seeing over and over again is that the cost of inference is

72
00:04:43.620 --> 00:04:44.453
exploding,

73
00:04:44.620 --> 00:04:49.460
driven by increasing the amount of context we put into agents and the number of

74
00:04:49.560 --> 00:04:51.600
loops and iterations we do around it.

75
00:04:52.180 --> 00:04:56.860
So here you can see like the average input tokens going up over time and

76
00:04:57.220 --> 00:05:01.640
average output tokens when across all API requests on

77
00:05:01.740 --> 00:05:06.420
OpenRouter.
And if you just look at this on a session basis,

78
00:05:06.480 --> 00:05:07.620
it gets even wilder.

79
00:05:08.100 --> 00:05:11.580
These are just like per request average tokens.

80
00:05:12.920 --> 00:05:16.780
So what we want to do is get us back to value-based pricing.

81
00:05:16.860 --> 00:05:18.960
Rather than starting with your pricing models,

82
00:05:19.020 --> 00:05:22.780
consider the way inference fits into the existing models you have.

83
00:05:22.840 --> 00:05:26.120
And here's a framework that I suggest you try.

84
00:05:27.780 --> 00:05:32.460
Classify the actual tasks that you are

85
00:05:32.820 --> 00:05:33.653
building around.

86
00:05:33.900 --> 00:05:38.800
So rather than simply throwing a ton of data at GPT-5.5

87
00:05:39.040 --> 00:05:43.680
or the latest or whatever frontier model you want and asking for multiple

88
00:05:44.100 --> 00:05:48.640
outputs all in the same request, break the problem down into components.

89
00:05:50.280 --> 00:05:55.180
Then map those components, map those subtasks to optimal models.

90
00:05:55.240 --> 00:05:58.400
And I'll show a couple examples of how people are doing this successfully in the

91
00:05:58.460 --> 00:06:03.160
industry very soon. So we mapped out all the inference.

92
00:06:03.400 --> 00:06:07.960
We wrote a report called State of AI at the end of last year in collaboration

93
00:06:08.000 --> 00:06:09.880
with Andreessen Horowitz, and

94
00:06:11.560 --> 00:06:15.160
we compiled a bunch of data about the different use cases across the market,

95
00:06:15.720 --> 00:06:19.060
how much it costs on average to serve those use cases,

96
00:06:19.600 --> 00:06:23.960
and the amount of tokens that we typically see associated with it.
So

97
00:06:25.220 --> 00:06:27.380
we mapped out all inference by the category of usage,

98
00:06:27.800 --> 00:06:32.270
and you can see the highest token volumes are on top,

99
00:06:32.430 --> 00:06:34.270
and it gets more expensive over to the right.

100
00:06:35.990 --> 00:06:40.670
What we noticed is that tasks are breaking down on a type of

101
00:06:40.750 --> 00:06:44.990
determinism with a heavy skew towards more variable workloads.

102
00:06:45.430 --> 00:06:49.930
So tasks over in this area are deterministic.

103
00:06:50.490 --> 00:06:52.430
Some people call them saturated tasks.

104
00:06:53.770 --> 00:06:57.550
They're ones like looking for a certain number,

105
00:06:58.150 --> 00:07:00.590
like is this PR HIPAA compliant,

106
00:07:01.070 --> 00:07:04.650
or can you classify this input as needing

107
00:07:05.710 --> 00:07:06.750
moderation or not?

108
00:07:08.790 --> 00:07:11.490
Or can you just translate this input into another language?

109
00:07:11.850 --> 00:07:16.230
Can you check to see if there's some kind of syntax error here?

110
00:07:16.870 --> 00:07:20.530
These really bespoke deterministic tasks, if you break it down far enough,

111
00:07:21.230 --> 00:07:22.690
you can optimize really aggressively.

112
00:07:23.890 --> 00:07:28.650
There's also variable tasks where you have no idea what the output is going to

113
00:07:28.690 --> 00:07:33.350
be. Like your users might be doing wild auto research pipelines.

114
00:07:33.410 --> 00:07:36.190
They might be like exploring their next novel.

115
00:07:36.250 --> 00:07:40.490
You have no idea what it's going to be. Then you end up spending more,

116
00:07:40.570 --> 00:07:45.150
because you just can't optimize that level of unknown.

117
00:07:46.310 --> 00:07:49.010
But if you can break your problem down, you can do quite a bit.

118
00:07:49.610 --> 00:07:52.610
And then there's an area in the middle, which we call semivariable,

119
00:07:53.010 --> 00:07:54.950
where we kind of know what we want,

120
00:07:55.310 --> 00:07:57.810
but the outputs could vary pretty dramatically.

121
00:07:57.870 --> 00:08:01.370
And this would be like a general poll request review, for example,

122
00:08:01.830 --> 00:08:05.530
or trivia, or completing the chapter of a novel,

123
00:08:07.170 --> 00:08:09.830
even doing like kind of open-ended legal requests.

124
00:08:11.330 --> 00:08:13.010
So you might be skeptical about this.

125
00:08:13.390 --> 00:08:17.550
Why not just use the best intelligence all the time and find ways to charge for

126
00:08:17.590 --> 00:08:21.430
it? Isn't that what the best companies are doing? Like if you're serious,

127
00:08:21.490 --> 00:08:23.370
shouldn't you just use frontier models?

128
00:08:25.710 --> 00:08:29.210
But this is not what we actually observe in the ecosystem,

129
00:08:29.450 --> 00:08:34.390
that the most serious teams are actually heavily optimizing their model usage.

130
00:08:34.750 --> 00:08:38.590
And that's because they realize they can do a lot more by cutting costs,

131
00:08:38.650 --> 00:08:43.230
and the costs are not insubstantial at all. So from a volume perspective,

132
00:08:43.910 --> 00:08:47.690
for months, this hasn't been true. When you look at popular models,

133
00:08:48.240 --> 00:08:53.160
we see that the lowest costs overtake the highest cost models

134
00:08:53.440 --> 00:08:58.440
in token volume. And we exclude free data here.

135
00:08:58.760 --> 00:09:03.760
So this is cheapest versus most expensive popular models by token volume.

136
00:09:04.040 --> 00:09:08.090
And you can see that like the lower cost models are actually growing faster.

137
00:09:09.240 --> 00:09:10.073
If you start,

138
00:09:10.160 --> 00:09:14.280
we start in Q4 with the inflection in open source models and lightweight models,

139
00:09:14.440 --> 00:09:16.040
including those from Gemini,

140
00:09:16.480 --> 00:09:19.400
these models have been like surging in usage recently because people are

141
00:09:19.520 --> 00:09:21.030
figuring out how to offload tasks.

142
00:09:21.360 --> 00:09:26.200
People are breaking down their larger tasks into subtasks.
And what

143
00:09:26.400 --> 00:09:28.280
about programming, which is pretty open-ended.

144
00:09:29.670 --> 00:09:31.040
This is a really hard one to make deterministic.

145
00:09:32.080 --> 00:09:36.000
Surely this is only going to land on top-tier models, right?

146
00:09:37.400 --> 00:09:39.960
Even programming, I would challenge you to decompose it.

147
00:09:40.640 --> 00:09:44.840
Like I mentioned before, you can put code review as a semivariable task,

148
00:09:45.160 --> 00:09:49.520
and we see a lot of teams using nonfrontier models,

149
00:09:49.680 --> 00:09:54.160
but with very clear guidelines about what to check each poll request for when

150
00:09:54.200 --> 00:09:58.960
they make their own PR review bots. For code scanning,

151
00:09:59.200 --> 00:10:03.720
for just looking for simple accessibility violations in

152
00:10:04.190 --> 00:10:08.080
HTML, these can be done in a really cheap bespoke way.

153
00:10:09.400 --> 00:10:14.200
And we've seen this shift happen now across the whole year.

154
00:10:14.360 --> 00:10:16.720
So you can see the top five most expensive models,

155
00:10:16.920 --> 00:10:20.440
and then you can see everything else taking up token share.

156
00:10:20.600 --> 00:10:24.240
And the top five most expensive models are still dominating by dollar volume,

157
00:10:24.560 --> 00:10:28.600
but tokens give you a sense of the tasks and time

158
00:10:29.550 --> 00:10:30.450
that models are being spent,

159
00:10:32.040 --> 00:10:34.250
that people are spending doing inference on different models.

160
00:10:36.600 --> 00:10:39.430
This is kind of another way of visualizing the landscape here.

161
00:10:39.790 --> 00:10:44.360
You can see open source in green on the left and closed source on the right,

162
00:10:45.590 --> 00:10:50.550
where closed source in general is a higher cost per million

163
00:10:50.590 --> 00:10:51.423
tokens.

164
00:10:51.630 --> 00:10:56.450
And then you can see kind of frontier models down below where

165
00:10:57.450 --> 00:11:01.670
usage gets... It is actually lower than the open source models today.

166
00:11:03.650 --> 00:11:06.330
So look into your tasks, look at which models they're using,

167
00:11:06.350 --> 00:11:07.470
and look at how much you're paying.

168
00:11:07.990 --> 00:11:10.870
Let me give a concrete example of a company that did this really well.

169
00:11:11.550 --> 00:11:15.970
Shopify had a single-purpose agent whose purpose was to crawl and

170
00:11:16.190 --> 00:11:20.190
analyze shops for different kinds of outcomes.

171
00:11:21.150 --> 00:11:23.990
And it started out as a one-shot agent

172
00:11:25.850 --> 00:11:29.210
using a GPT series model.
And what they did was,

173
00:11:29.230 --> 00:11:33.790
they would like throw tons of context at GPT-5 and then

174
00:11:34.130 --> 00:11:36.850
ask for a couple different outcomes like,

175
00:11:37.650 --> 00:11:40.010
is there like a fraud situation

176
00:11:40.110 --> 00:11:44.930
here? Can agents crawl this site?

177
00:11:45.150 --> 00:11:48.710
They were asking a couple different questions about the site all in a single

178
00:11:48.750 --> 00:11:53.750
prompt. This is a perfect example of a good task to break down and

179
00:11:53.810 --> 00:11:55.590
do in multiple requests instead of one.

180
00:11:56.450 --> 00:12:00.730
So what they did was they broke up their usage into three separate

181
00:12:00.770 --> 00:12:01.910
subtasks and

182
00:12:04.120 --> 00:12:06.850
they were able to lower the spend from,

183
00:12:07.250 --> 00:12:11.970
I think it was like $5.5 million to about $75,000 a year.

184
00:12:12.720 --> 00:12:17.410
And it improved results like the F1 score,

185
00:12:17.810 --> 00:12:20.490
which is kind of a harmonic mean between precision and recall,

186
00:12:20.830 --> 00:12:25.830
improved for their compound agent relative to the one-shot GPT-5

187
00:12:26.270 --> 00:12:30.490
agent. So these are like the three separate subagents they ended up developing,

188
00:12:30.830 --> 00:12:33.710
and they just would just run the subagents in parallel.

189
00:12:35.930 --> 00:12:39.470
You don't need to be at their scale though to see this level of savings.

190
00:12:39.550 --> 00:12:43.950
We see a surprisingly persistent premium for the best intelligence versus

191
00:12:43.990 --> 00:12:48.610
acceptable intelligence. This is one of those crazy charts with a crazy y-axis,

192
00:12:49.150 --> 00:12:52.210
but like here you get the priciest 10 models,

193
00:12:52.430 --> 00:12:54.710
here you get the cheapest 10 models on OpenRouter,

194
00:12:55.730 --> 00:12:57.410
and here you get a lot of time.

195
00:12:58.150 --> 00:13:02.250
And what we see is that the delta between the priciest 10 models and the

196
00:13:02.310 --> 00:13:04.630
cheapest 10 models, this is a pretty interesting chart,

197
00:13:05.370 --> 00:13:07.870
is reliably 25x.

198
00:13:07.990 --> 00:13:11.690
I don't know if anybody figured that out from the y-axis, but yeah,

199
00:13:11.750 --> 00:13:15.090
that's like a 25x delta. So that's basically

200
00:13:17.070 --> 00:13:21.810
what the market saves when moving to lower cost models.

201
00:13:23.830 --> 00:13:28.330
The dropping cost of intelligence also means you can do a lot more with these

202
00:13:28.390 --> 00:13:30.590
lower cost models, and the pace is picking up.

203
00:13:30.910 --> 00:13:35.730
So here's a chart from Artificial Analysis where you can see this

204
00:13:36.170 --> 00:13:39.290
applying across all intelligence levels. Here,

205
00:13:39.570 --> 00:13:42.910
this blue line is kind of a lower intelligence level,

206
00:13:43.570 --> 00:13:48.030
and you can see that we started at this price point and after about four months,

207
00:13:48.450 --> 00:13:53.070
it dropped, it dropped, it dropped again. The next intelligence level,

208
00:13:53.610 --> 00:13:54.670
same pattern happened.

209
00:13:54.730 --> 00:13:58.730
It just takes a couple months after each intelligence level is achieved,

210
00:13:58.850 --> 00:14:02.640
which is basically each leap forward in

211
00:14:04.630 --> 00:14:07.150
LLMs for cost drop.

212
00:14:07.630 --> 00:14:12.530
And so you can kind of like reliably expect declining costs and to get down to

213
00:14:12.570 --> 00:14:14.030
this 25x savings level.

214
00:14:15.370 --> 00:14:19.810
We did some research on this last year, like I mentioned,

215
00:14:20.170 --> 00:14:24.230
and we noticed a striking effect in how models are adopted as a result.
People

216
00:14:24.290 --> 00:14:26.490
basically try models really quickly,

217
00:14:27.170 --> 00:14:32.110
and there's always this power user cohort that finds

218
00:14:32.490 --> 00:14:35.270
an amazing use case for a new model, and they're like, "Whoa,

219
00:14:35.610 --> 00:14:38.110
this solves all my issues with all the previous models.

220
00:14:38.170 --> 00:14:40.410
I don't have to triple prompt anymore,

221
00:14:40.790 --> 00:14:44.230
or I finally don't have to create these guardrails that I used to have in my

222
00:14:44.290 --> 00:14:46.640
code," and they stick to it. And

223
00:14:48.550 --> 00:14:52.950
you'd expect that if this is true, then we should be able to find a cohort,

224
00:14:53.390 --> 00:14:58.190
like a first-week cohort that has much higher retention than all other cohorts.

225
00:14:58.250 --> 00:15:02.790
These are the power users who discover these latent capabilities in new model

226
00:15:02.810 --> 00:15:03.643
launches.

227
00:15:03.690 --> 00:15:08.230
So we call this the "glass slipper effect," and you can illustrate it here.

228
00:15:08.370 --> 00:15:13.190
So this is a chart of Claude for Sonnet adoption,

229
00:15:13.650 --> 00:15:17.310
and you can see that the first cohort, which is in orange here,

230
00:15:19.050 --> 00:15:23.610
this is the retention of people who adopt it the month that it launches,

231
00:15:24.170 --> 00:15:29.110
and it's significantly bigger than the retention of cohorts who adopt

232
00:15:29.150 --> 00:15:32.570
it in subsequent months.
So this is the glass slipper effect in action.

233
00:15:32.630 --> 00:15:37.450
It's the fact that there are these tons of enthusiasts and tons of companies who

234
00:15:37.510 --> 00:15:42.430
are actively searching for models that work for new use cases

235
00:15:42.770 --> 00:15:44.910
where they're struggling with the existing frontier.

236
00:15:47.130 --> 00:15:51.210
So I believe that you can try on these glass slippers in a sense,

237
00:15:51.790 --> 00:15:54.150
and that your path to getting inference spend under control

238
00:15:55.730 --> 00:15:58.010
will allow you to return to value-based pricing,

239
00:15:58.470 --> 00:16:02.230
the value-based pricing that you intend, when you launch your pricing plan.

240
00:16:02.510 --> 00:16:04.650
And if you can get your costs down quite a bit,

241
00:16:04.990 --> 00:16:06.850
then you can just charge the way you want,

242
00:16:07.450 --> 00:16:09.290
which is just the way your customers want,

243
00:16:09.890 --> 00:16:12.730
the way your customers actually see your service.

244
00:16:15.110 --> 00:16:18.630
This is the thought experiment I've been giving everyone:

245
00:16:18.870 --> 00:16:23.050
whether you're planning your personal agent or the spend for your company,

246
00:16:23.890 --> 00:16:28.050
you should analyze your tasks by how deterministic they are and start testing

247
00:16:28.090 --> 00:16:29.450
them against more models.

248
00:16:31.810 --> 00:16:36.410
Another way to look at this is that you can make the token disappear from your

249
00:16:36.450 --> 00:16:40.270
calculations and turn it into an infrastructure decision,

250
00:16:40.650 --> 00:16:44.240
so that you can reclaim your value-based pricing. And

251
00:16:45.850 --> 00:16:49.490
this is an incredible way to get started. One way,

252
00:16:50.610 --> 00:16:55.070
the way that Shopify did, is by using a framework like DSPy,

253
00:16:55.850 --> 00:16:57.430
which works really well with OpenRouter.

254
00:16:57.630 --> 00:17:02.170
You can compile your code against a prompt that it'll

255
00:17:02.330 --> 00:17:03.670
automatically optimize.

256
00:17:04.310 --> 00:17:09.190
And the moment you want to try out a new model like Qwen or a new open

257
00:17:09.250 --> 00:17:13.090
source model that launches, you just change the model slug, recompile,

258
00:17:13.530 --> 00:17:16.470
and find the new prompt that's going to work for your use case.

259
00:17:18.350 --> 00:17:21.170
Okay. So the question is,

260
00:17:21.170 --> 00:17:25.550
"Do you see availability and capacity being phased out for older models?

261
00:17:25.750 --> 00:17:26.970
We find our glass slipper;

262
00:17:26.970 --> 00:17:30.250
can we rely on them staying available?" So this is a good question.

263
00:17:30.650 --> 00:17:35.570
We frequently see models get deprecated across the ecosystem.

264
00:17:36.050 --> 00:17:40.530
It is amazing how many people are still using Gemini 2.5 Flash.

265
00:17:42.470 --> 00:17:47.170
There are a lot of open source models that get dropped from a lot of the open

266
00:17:47.230 --> 00:17:50.510
source inference providers. And this goes for all inference providers,

267
00:17:50.650 --> 00:17:52.690
hyperscalers, neoclouds,

268
00:17:52.990 --> 00:17:57.610
they're all dropping models when they find that demand on their platform just

269
00:17:57.670 --> 00:18:01.990
does not justify the cost to keep the model running.

270
00:18:02.810 --> 00:18:06.510
So what do you do about deprecations? First,

271
00:18:06.790 --> 00:18:09.510
OpenRouter was partly made to help with this.

272
00:18:10.370 --> 00:18:12.770
We aggregate the whole market for you,

273
00:18:13.270 --> 00:18:17.030
so that when a model gets deprecated on one provider, your app stays up.

274
00:18:18.430 --> 00:18:19.263
Moreover,

275
00:18:19.630 --> 00:18:24.490
you can provide fallback models or just use the Auto Router so that if the whole

276
00:18:24.590 --> 00:18:28.010
model were to disappear-which means all providers on the market disappear,

277
00:18:28.710 --> 00:18:29.543
your app stays up.

278
00:18:30.250 --> 00:18:35.130
And we're seeing a shift towards auto routers, partly for this reason,

279
00:18:35.990 --> 00:18:40.610
but also a shift towards aggregating inference.
We

280
00:18:40.710 --> 00:18:44.680
started back in, I think, June of 2023,

281
00:18:47.430 --> 00:18:50.110
we put all providers on a model page,

282
00:18:50.410 --> 00:18:52.650
so that we could increase everyone's uptime. And that way,

283
00:18:52.670 --> 00:18:53.850
when you use OpenRouter,

284
00:18:54.210 --> 00:18:58.950
you get the combination of uptime from all providers in the

285
00:18:59.030 --> 00:18:59.863
ecosystem,

286
00:19:00.350 --> 00:19:05.030
and you also get flexibility on which features you want. If you want speed,

287
00:19:05.210 --> 00:19:06.550
if you want certain samplers,

288
00:19:07.210 --> 00:19:10.350
we can aggregate the features of the ecosystem for you,

289
00:19:10.890 --> 00:19:14.470
and that's a big reason people use us-especially for inferencing open source

290
00:19:14.510 --> 00:19:17.730
models. So when old providers drop out,

291
00:19:17.850 --> 00:19:19.300
you can continue using them on OpenRouter.

292
00:19:19.670 --> 00:19:22.140
And then if you're worried about the model disappearing,

293
00:19:23.160 --> 00:19:25.800
I just urge kind of moving to routers in general.

294
00:19:25.970 --> 00:19:28.020
We allow you to define your own routering policies.

295
00:19:28.200 --> 00:19:32.920
And then we also build routers that help you target models just

296
00:19:32.990 --> 00:19:37.880
above a certain intelligence level-a new one that we're about to announce soon

297
00:19:37.990 --> 00:19:41.140
that is available for people today, called the Pareto Router.

298
00:19:43.180 --> 00:19:47.420
The second question is, "How close to frontier do you think the new DeepSeek,

299
00:19:47.520 --> 00:19:49.120
Kimmy, and Qwen models actually are?

300
00:19:49.380 --> 00:19:53.160
Can we be more aggressive in moving these highly variable workloads yet?"

301
00:19:54.730 --> 00:19:56.940
So there are

302
00:20:00.280 --> 00:20:04.660
a couple nuances to this. First, I think, like I mentioned before,

303
00:20:04.710 --> 00:20:06.340
when you optimize your inference,

304
00:20:06.840 --> 00:20:11.540
you can get better than frontier quality out of all of these

305
00:20:11.560 --> 00:20:12.393
models.

306
00:20:12.730 --> 00:20:17.600
That Shopify example that I mentioned moved to Qwen from GPT-5,

307
00:20:20.470 --> 00:20:22.260
but it requires optimization,

308
00:20:22.340 --> 00:20:27.230
requires you to decompose your task and to be careful about what your

309
00:20:27.260 --> 00:20:31.640
prompt is. But the cost savings are quite worthwhile,

310
00:20:31.640 --> 00:20:35.400
and you can improve performance on your agent just

311
00:20:35.420 --> 00:20:37.490
by putting the effort into optimizing it.

312
00:20:40.710 --> 00:20:45.260
The nuance here is that there are just a lot of nondeterministic use

313
00:20:45.300 --> 00:20:47.040
cases in major apps today.

314
00:20:47.420 --> 00:20:49.880
And if you're building a general purpose coding agent,

315
00:20:50.860 --> 00:20:53.640
there's just not that much task decomposition out there.

316
00:20:53.960 --> 00:20:55.800
You can label the chats,

317
00:20:58.560 --> 00:21:02.820
you can have a CI agent that runs separately from the code generation agent.

318
00:21:03.160 --> 00:21:08.000
You can do live feedback to the model as it's outputting code.

319
00:21:08.440 --> 00:21:10.740
So there are some ways of breaking things down, but

320
00:21:12.280 --> 00:21:14.920
a lot of the aspects of AI today are about,

321
00:21:15.260 --> 00:21:18.340
how do we deal with this emerging market where we don't even know what to build?

322
00:21:19.260 --> 00:21:22.380
And that's where we see frontier models still be really,

323
00:21:22.440 --> 00:21:27.380
really strong.
"What do you mean 'make the token disappear'?" So what I mean

324
00:21:27.400 --> 00:21:32.280
by that is, if you are really worried about tokens,

325
00:21:34.280 --> 00:21:37.540
it generally means that you have not optimized your app.

326
00:21:37.840 --> 00:21:42.740
You haven't broken the task down into subtasks.

327
00:21:42.740 --> 00:21:44.200
And so when I say,

328
00:21:44.200 --> 00:21:47.940
"Make the token disappear," I just mean like think about the actual tasks that

329
00:21:47.980 --> 00:21:50.520
you're accomplishing and how much they cost on average,

330
00:21:51.080 --> 00:21:54.200
rather than the total number of tokens you're consuming in the background.

331
00:21:57.520 --> 00:22:02.240
"Will OpenRouter accept stablecoin payment natively?" So

332
00:22:02.840 --> 00:22:05.800
we've been accepting stablecoin payments since May of 2023.

333
00:22:09.240 --> 00:22:11.340
I think we were one of the earliest platforms to do it,

334
00:22:11.880 --> 00:22:16.840
and we also built a programmatic way for agents to do it prior to x402.

335
00:22:17.840 --> 00:22:22.780
Since then, protocols like x402 and MPP have emerged that do this

336
00:22:23.620 --> 00:22:24.660
in a better way,

337
00:22:25.180 --> 00:22:29.900
so expect to see some more work from us soon to make stablecoin payments even

338
00:22:30.000 --> 00:22:34.020
easier for agents.
"Do you have any insight between the time from frontier model

339
00:22:34.060 --> 00:22:38.400
intelligence or capabilities and open source models?"

340
00:22:41.720 --> 00:22:43.440
We do work with the Model Labs,

341
00:22:45.380 --> 00:22:50.220
but I wouldn't say we have any insights that we can share that are not

342
00:22:50.300 --> 00:22:52.220
public. The

343
00:22:53.880 --> 00:22:58.700
frontier model intelligence generally leads open source

344
00:22:58.780 --> 00:23:02.680
model intelligence by about six months.

345
00:23:02.680 --> 00:23:05.620
I'd say that when a new open source model launches,

346
00:23:06.200 --> 00:23:08.480
there's typically a ton of hype on Twitter,

347
00:23:09.020 --> 00:23:13.580
and be careful how assuaged you get by the hype the week the model

348
00:23:13.600 --> 00:23:14.433
launches.

349
00:23:15.460 --> 00:23:19.280
It's really important to be scientific about whether it's actually working.

350
00:23:19.560 --> 00:23:21.460
What I see a lot of people do, for example,

351
00:23:21.520 --> 00:23:24.240
is a new model will launch or a new harness will launch,

352
00:23:24.680 --> 00:23:28.340
and then they'll give it like three prompts manually and they'll be like,

353
00:23:28.340 --> 00:23:32.300
"I don't think it's better. I think it's worse." Or they'll give it one,

354
00:23:32.420 --> 00:23:34.060
two prompts in Claude Code and be like, "Oh my God, I think it's better.

355
00:23:34.060 --> 00:23:37.580
It's better." And then they'll post about it.

356
00:23:37.900 --> 00:23:42.660
And we're just in this phase of the market where a sample

357
00:23:42.700 --> 00:23:45.320
size of one is okay, which is crazy.

358
00:23:46.020 --> 00:23:50.780
So increase your sample size and do actual studies because

359
00:23:51.020 --> 00:23:54.400
these things are nondeterministic black boxes,

360
00:23:54.780 --> 00:23:59.220
and you have to learn from real data in the market in addition to

361
00:23:59.800 --> 00:24:02.720
running this on evals that you write yourself.

362
00:24:06.640 --> 00:24:09.860
I think that's the last one. Oh, "A year ago,

363
00:24:09.860 --> 00:24:09.860
I asked you if users wanted to pay per use instead of sign up and buy credits.

364
00:24:09.860 --> 00:24:10.613
You said 'No.' Has that changed in the last year?"

365
00:24:23.260 --> 00:24:26.420
There are still pay-per-use use cases.

366
00:24:26.920 --> 00:24:31.400
It has not changed that it's a huge need.

367
00:24:32.880 --> 00:24:37.220
I mean, I would say that the pay per use is really,

368
00:24:37.280 --> 00:24:42.220
really powerful when you have agents being built that have no idea how they're

369
00:24:42.240 --> 00:24:43.073
going to be used,

370
00:24:43.460 --> 00:24:46.220
because they have no idea what services they're going to need as a result.

371
00:24:46.940 --> 00:24:50.700
And OpenClaw actually was a really big moment for us,

372
00:24:52.600 --> 00:24:57.360
for this case when OpenClaw started, it had these two actions,

373
00:24:57.460 --> 00:24:59.820
a heartbeat and everything else.

374
00:25:00.280 --> 00:25:05.220
And a lot of the setup tasks from OpenClaw just don't work unless you're using a

375
00:25:05.360 --> 00:25:06.200
really good model,

376
00:25:06.340 --> 00:25:10.960
but then those really good models would charge you a request every

377
00:25:11.060 --> 00:25:15.220
5, 10 minutes for the heartbeat.
So people really wanted to move to the auto

378
00:25:15.300 --> 00:25:19.120
router, and that's how our Auto Router started getting a lot of growth.

379
00:25:19.660 --> 00:25:24.180
So these pay-per-use use cases tend to

380
00:25:24.240 --> 00:25:28.800
emerge when some sort of general purpose app like OpenClaw needs some

381
00:25:28.880 --> 00:25:33.380
service that the original developer never knew would be needed.

382
00:25:34.100 --> 00:25:36.700
And I think we'll see more in the coming 6

383
00:25:39.660 --> 00:25:43.580
to 12 months that does that. Okay.

384
00:25:44.880 --> 00:25:45.580
Thanks for having me.

