We’re having great results from our Chatbase chat engine on fintechbenchmark.com. But we have to look ahead, and smaller models have caught our attention.
For years, the industry assumed bigger was always better, with companies building increasingly massive data centres to house trillion-parameter models. But there’s a shift happening toward more specialized, efficient models – and for good reason.
The Power of Specialization
If you’re using AI to provide intelligent search of a website with 4,000 fintech products, you don’t need the model trained on Shakespeare, quantum physics, or the Treaty of Westphalia. The types of questions are limited and the subject matter is bounded, so a smaller, specialized model is not just feasible – it’s often superior.
Don’t underestimate the importance of this. Large general-purpose models remain crucial for broad AI applications, but there are far more cases in limited domains where specialized models deliver better results at a fraction of the cost.
Understanding Model Size
The size and complexity of AI models is measured in “parameters” – think of these as the knobs and dials the AI adjusts during training to recognize patterns. A 7 billion parameter model has 7 billion tiny settings tuned by analyzing text. Large-scale models like GPT-4 have around two trillion parameters, making 7 billion genuinely small by comparison. (If you want to know more, search for “neural networks”.)
Here’s the remarkable part: a 7 billion parameter model needs only about 28GB of storage and can run on high-end computer hardware. You don’t need a data centre.
The Infrastructure Question
This opens up interesting options, but also some challenges:
- Can’t use standard web hosting: Our VPS doesn’t have the storage, RAM, or processing power. You’d need a dedicated server with GPU (Graphics Processing Unit) chips.
- Specialist GPU hosting: Services like Replicate, Modal, or RunPod offer pay-per-use GPU access, or you could use cloud platforms like AWS.
- When it makes sense: Self-hosting only pencils out for high volumes or where speed is critical (think fraud detection or trading applications). For most applications, API services remain more cost-effective.
- Security advantage: In-house hosting gives you complete control over data – important for fintech compliance requirements.
Why GPUs?
Graphics Processing Units were originally developed for rendering graphics quickly, but turned out to be excellent at the parallel processing AI requires. Now there are AI-specific chips under development (from companies like Groq, Cerebras, and others) that will be even faster and more power-efficient than GPUs.
The Future is Hybrid
We’re likely heading toward a world where your phone runs a local AI model for routine tasks but outsources complex questions to cloud services. It’s not a question of if, but when.
For fintech companies evaluating AI infrastructure, the key takeaway is this: match your model to your problem. Don’t pay for trillion-parameter generalist models when a lean, specialized 7 billion parameter model trained on your domain will deliver better, faster, cheaper results.
This area is still in its infancy. Watch this space for major developments.
Leave a comment