Kordane ForgeModel evaluation & optimisation

Match the model to the job.

Make model choice an engineering decision. Compare quality, latency and cost against your task, then find the smallest permitted model that meets your requirements.

Training and evaluation available in pilots · selection, runtime and distributed evaluation planned

Start with your task.

Define what a good result looks like. Evaluate candidates on relevant examples and acceptance criteria.

Make efficiency measurable.

Compare smaller, specialised and frontier models on the same work. Keep the trade-offs visible.

Keep the reason for the choice.

Record the selected candidate, the alternatives and the conditions for fallback. Model optimisation stays inside your information policy.

Explore the trade-offs

Change the task.
See the decision.

Quality, cost and deployment requirements change the choice. Explore an illustrative selection, with information policy setting the permitted options.

Synthetic decision logic.Illustrative rules, not measured benchmarks.

Deterministic rules show how BoundSense (permission and environment) and Forge (model sufficiency) decide together. Nothing you select is transmitted.

Decision Small local model

Task classification
·
Information class
·
Allowed environment
·
Selected model class
·
Fallback route
·
External transfer
·
Decision reason
·
Evidence status
·
Explore private distributed evaluation

Product roadmap

Forge Private Distributed Evaluation

PlannedPlanned architecture. Not implemented, and not available in pilots today.

Evaluate permitted candidate models across private or isolated datasets without centralising their raw evaluation examples. BoundSense-authorised nodes control what runs locally and what metrics or evidence may leave. This is distributed evaluation, not federated training.

  1. Forge creates the evaluation package

    One immutable package identifies the task, the candidate, the runner and the metrics, so every node evaluates the same thing and the run can be repeated.

  2. BoundSense authorises the package at each node

    Each node authorises independently: which package, which candidate, which local dataset reference, and what may leave. One node’s permission never extends to another.

  3. Evaluation executes locally

    Each environment runs the candidate against its own dataset, under its own access controls, and records complete evidence locally.

  4. Raw evaluation examples stay in place

    The examples themselves are not centralised in order to score them. They remain inside the approved environment that holds them.

    Raw examples contained locally

  5. Only policy-authorised evidence leaves

    Each node emits the metrics, findings and evidence its local policy explicitly allows, and nothing else.

    Permitted to leaveMetrics · safety findings · failure counts · local evidence identities

  6. Forge combines the results

    Aggregation records which nodes participated, which were unavailable and which were excluded, so a partial run is never read as a complete one.

  7. Global and per-node thresholds are evaluated separately

    Global, per-node, subgroup and jurisdiction criteria are checked in their own right.

    Global threshold: passed

  8. A required node fails

    One required environment falls below its threshold. The global average still passes. That is exactly the case the design exists for.

    Required node threshold: failed

  9. Promotion is held

    A passing global metric does not release promotion on its own. The candidate is held until the required node passes or the task specification changes.

    Promotion: held

  10. The combined record keeps the failure visible

    One record links the aggregate result to each node’s local evidence identity. Failed and refused evaluations are not rewritten to make a candidate look successful.

Distributed evaluation across private datasets. This roadmap describes evaluation, not federated training.

BoundSense-authorised nodesPlanned

The same control plane, placed in each customer-controlled environment rather than only in front of one.

  • What may enter. Which evaluation package or query is allowed to reach this environment at all.
  • What may execute. Which candidate model may run locally, from the set already permitted for the task.
  • What may participate. Which local dataset or source is eligible for this specific run.
  • What may leave. Which metrics, excerpts, structured facts or local answers policy authorises out of the node.
  • What record remains. The crossing record each node keeps locally, and the identity it contributes to the combined record.

A node is a placement of BoundSense, not a separate product. Nothing here adds a product family or a navigation entry.

See this scenario in the interactive system diagram

A first conversation

Start with
one workflow.

Tell us the task and where the information must stay. We'll show a relevant example and discuss whether a scoped pilot makes sense.

Request a demo