Why Small Local Models Matter for Enterprise Automation
Small local models can handle recurring, bounded tasks where cost, privacy, latency, and operational control matter more than general-purpose brilliance.
Research blog
Seeded essays on model evaluation, RAG, prompt routing, local infrastructure, and the bridge from classical computer vision to modern AI.
Small local models can handle recurring, bounded tasks where cost, privacy, latency, and operational control matter more than general-purpose brilliance.
A preliminary look at model size, quantisation, context settings, and inference behavior as interacting variables rather than isolated knobs.
Prompt routing turns AI adoption into architecture: classify the request, select the right model, enforce policy, and preserve auditability.
Practical RAG depends on retrieval quality, lexical search, vector search, reranking, and answer constraints, not just embeddings.
A field note on LM Studio, Ollama, llama.cpp, vLLM exploration, and the tradeoffs of running inference locally.
A line from 2012 HAAR and QR experiments to modern multimodal systems: the goal remains turning perception into useful action.