Model evaluation is task-specific
Benchmarks measure the evaluated scope only; they do not establish general performance.
Small-model success is not general capability
A small model clearing one task threshold implies nothing about other tasks.
Benchmark leakage and dataset defects
Training and evaluation can inherit or amplify dataset defects and leakage; data quality remains a shared responsibility.
Local execution is not security
Running locally reduces egress exposure but does not by itself make a system secure.
Fallback routes can increase exposure
Escalating to a larger external model widens the crossing; fallback paths need the same policy scrutiny as primary routes.
Channel connectors depend on platforms
Axon conversation connectors depend on external platform controls and availability.
Human agents are a boundary
People handling escalated conversations remain part of the trust boundary.
Distributed execution is not a privacy technology
Keeping raw examples and corpora in place narrows what crosses. It is not secure aggregation, differential privacy, confidential computing or zero-knowledge protection, and the metrics or excerpts that do cross can still carry inference risk.
Aggregate results can hide local failures
A global metric can pass while a required environment fails. Distributed evaluation is designed to keep that failure visible and hold promotion; an aggregate on its own is not a release decision.
Distributed capabilities are planned
Private distributed evaluation and governed distributed retrieval are roadmap architecture. Neither is implemented, and neither is available in pilots today.