AI Models
Local LLMs, vision-language, multimodal and embedding models, quantization, model switching, and multiple models running simultaneously.
Which models are actually useful on this hardware?
Real-World AI Infrastructure
Testing local AI systems, models and agents against real business workloads to understand what actually works in deployment.
Hardware. Models. Agents. Vision. Automation. Operations.
Real Applications • Real Workflows • Local AI • Agentic Systems • Business Deployment
Independent Evaluation Program
AI is moving beyond cloud-only experimentation. Compact AI workstations, edge AI systems, local AI servers and specialized AI appliances are becoming capable of running increasingly sophisticated models and autonomous agents locally.
The question is no longer simply "How powerful is the hardware?" The more useful question is "What can we actually deploy on it?"
RyanMatheson.ca is developing a practical evaluation program focused on answering that question. We test AI hardware and software against real-world workflows involving AI agents, computer vision, document processing, automation, QA, reporting and business operations.
The goal is to understand: what works, what doesn't, what works together, and where each platform makes sense.
What We Evaluate
Local LLMs, vision-language, multimodal and embedding models, quantization, model switching, and multiple models running simultaneously.
Which models are actually useful on this hardware?
Browser, coding and computer-use agents, tool calling, multi-step workflows, autonomous task execution, long-running jobs, and memory/context behavior.
Can the system run an agent that actually gets useful work done?
Image inspection, camera feeds, object detection, classification, OCR, document understanding, and real-time versus batch processing.
Can local AI understand what's happening in the physical environment?
Store inspections, compliance and safety workflows, 5S audits, document processing, PDF and report generation, corrective actions, inventory, and customer-service automation.
Can this system solve an actual business problem?
How the OS, drivers, GPU/NPU, AI runtime, Docker, Python, Node.js, Laravel, Ollama, llama.cpp, vLLM, TensorRT, ONNX, Playwright, MCP, databases and RAG systems work together.
How much work is required to turn the hardware into a usable platform?
A machine with excellent benchmark numbers may not be the best deployment platform. A less powerful machine may prove extremely effective for a specific business workload.
We're testing systems, not just specifications.
Deployment Models
The evaluation is intended to explore several realistic deployment models rather than a single idealized setup.
Local AI assistants, private document processing, internal knowledge assistants, inspection automation, camera and vision systems, customer-service and scheduling automation.
Data privacy, network isolation, local inference, internal document access, security, reliability, administration, and operating cost.
Systems deployed near the physical workload: retail, warehouses, manufacturing, field operations, cameras, sensors, and inspection stations.
A small business shouldn't need to understand GPUs, models, containers and inference servers. We explore which hardware and software combinations can become purpose-specific appliances for particular jobs.
Evaluation Methodology
Specifications tell us what a machine is capable of on paper. Real workloads tell us what it is capable of doing. The program therefore focuses on the complete system:
Hardware + Models + Runtime + Agents + Software + Workflow + Human Use
Compute capability, memory, GPU/NPU performance, storage, networking, expandability, physical footprint, power, thermal behavior, noise, setup complexity, driver maturity, container support, remote administration, maintenance and recovery.
Model compatibility, inference speed, vision performance, context capacity, concurrent and sustained workloads, agent performance, time-to-first-result, and throughput.
OS and driver/runtime maturity, Docker and container support, Python, Node.js and Laravel, Ollama, llama.cpp, vLLM, TensorRT, ONNX, Playwright, MCP, local databases and RAG systems.
What can it actually do? Who could use it? What problem does it solve? What replaces — or still depends on — cloud infrastructure? How much does deployment realistically cost? Can it run unattended and support multiple users as a reliable business appliance?
A Real Test Environment
One of the initial real-world environments available for evaluation is ComplyCam, an operations and compliance platform being developed and deployed by Ryan Matheson. Its workloads are one example of the type of work the evaluation environment can generate.
Automated browser interaction, application QA, inspection processing, image analysis, document generation, PDF reporting, workflow execution, error detection, agent-driven testing, and long-running automated jobs.
This demonstrates that the evaluation environment has an actual application and workload rather than relying solely on benchmark datasets.
Participation
AI hardware and infrastructure companies are welcome to participate in the evaluation program. We are interested in evaluating systems across a range of performance levels, architectures and intended use cases — from compact local AI workstations to edge systems and specialized AI appliances.
Companies participating in the program have an opportunity to demonstrate what their platform can do under practical workloads and to help establish where their technology fits in the emerging local-AI deployment landscape. Evaluation units, loaner systems, demonstration hardware, technical access and collaborative testing arrangements may be considered where appropriate.
Evaluations are intended to be practical and independent. Participation does not guarantee a positive result.
Evaluation Outputs
Hardware, software stack, AI capabilities, real-world workloads, deployment experience, strengths, limitations, and recommended use cases.
Configuration, models tested, workloads, performance observations, deployment architecture, and lessons learned.
Video demonstrations, screenshots, workflow demonstrations, agent runs, and business-use-case demonstrations.
Where does this system actually make sense — private local AI, small business AI, vision inspection, edge inference — and where it is less suitable.
Long-Term Vision
The local AI market is developing quickly. New AI workstations, edge systems, accelerators, models and agent platforms are appearing faster than traditional benchmarks can explain their practical value.
The goal of this program is to build a growing body of real-world evidence showing how these technologies perform when they are asked to do actual work. Over time, the evaluations can help businesses answer a much more useful question than "Which AI box is the fastest?"
"Which AI system is right for what we actually need to accomplish?"
Get Involved
If you're building local AI hardware, edge infrastructure, AI appliances, accelerators or deployment platforms, I'd be interested in discussing how your system could be evaluated against real workloads.