costs are in a large part tied to model parameter count, eg. a small model fits on one GPU, the biggest models require an entire rack
the primary difference you see through those prices is model size
that you don't seem much capability difference is why Big Ai is pushing for regulatory capture, because small models are nearly as good for many tasks at a fraction of the cost, undermining the business model they have used to justify the most expensive infra build out in human history
Also I think there are diminishing returns to larger models and the harness can make up for some weaknesses. Like maybe the huge model can hypothetically answer a question off the cuff but the small model can get some search results and think about those and give you an answer.
I am big believer (based on personal experience) that there is a ton of ROI in harness engineering, to go along with the context engineering, invariably intertwined
costs are in a large part tied to model parameter count, eg. a small model fits on one GPU, the biggest models require an entire rack
the primary difference you see through those prices is model size
that you don't seem much capability difference is why Big Ai is pushing for regulatory capture, because small models are nearly as good for many tasks at a fraction of the cost, undermining the business model they have used to justify the most expensive infra build out in human history
Also I think there are diminishing returns to larger models and the harness can make up for some weaknesses. Like maybe the huge model can hypothetically answer a question off the cuff but the small model can get some search results and think about those and give you an answer.
100%
I am big believer (based on personal experience) that there is a ton of ROI in harness engineering, to go along with the context engineering, invariably intertwined