Ontopix · Engineering Spec

Audibot v2
API Specification

A REST API for defining audit rubrics, submitting conversation batches, and retrieving defensible quality scores at population scale.

Document
SPEC-AUDIBOT-v2 · public
Version
2.0.0-draft.3
Date
2026-04-23
Owners
Audibot · Engineering
Authors
J. Roca, M. Esteban
Reviewers
Audibot Eng, Customer Trust, Legal
Status
Draft · for internal review
Audibot v2 API · Specification Contents 2

Contents

    1. 1.Overview3
    2. 2.API surface4
    3. 3.Data model5
    4. 4.Versioning & migration from v16
    SPEC-AUDIBOT-v2 · 2.0.0-draft.3 Ontopix · Confidential
    Audibot v2 API · Specification 1. Overview 3

    1. Overview

    Audibot evaluates customer-service conversations against a rubric the customer already uses with its human auditors. It replaces a 2 % sampling workflow with 100 % coverage, returning per-conversation scores with transcript-span evidence attached to every rating.

    The v2 API is a breaking revision of the v1 surface shipped in 2025-09. It introduces rubric versioning, streamed result delivery, and a normalized evidence schema. This document is the authoritative contract for integrators.

    1.1 Scope & non-goals

    In scope: the REST surface for submitting audits, retrieving results, and managing rubrics. Also in scope: the data contract, error taxonomy, idempotency rules, and migration from v1 endpoints.

    Out of scope: the LLM backbone configuration, scoring-model weights, batching optimization, and the internal worker architecture. These are implementation details and may change without a spec revision1.

    1.2 Compatibility with v1

    v2 is not backward-compatible with v1 at the wire level. Endpoints are served at /v2/*. The v1 surface at /v1/* continues to receive security fixes through 2026-09 and is then decommissioned.

    ⚠ Warning

    Rubric IDs are not portable between v1 and v2. A rubric defined via POST /v1/rubrics must be re-authored against the v2 schema; see §4 for the migration path.

    ◇ Rationale

    We chose a hard-break versioning strategy (separate path prefix, no graceful fallback) because v1's single-score rubric shape cannot represent v2's criterion-level evidence attachments. A translating shim layer would misreport uncertainty and was judged worse than a migration window.

    v2 also removes the synchronous /audits/evaluate endpoint. All audits are now asynchronous batches — a design change motivated by the latency variance observed when v1 was deployed against > 5k-conversation populations.

    SPEC-AUDIBOT-v2 · 2.0.0-draft.3 Ontopix · Confidential
    Audibot v2 API · Specification 2. API surface 4

    2. API surface

    2.1 Endpoint catalogue

    v2 endpoints. All responses are application/json unless noted.
    MethodPathPurposeIdempotent
    POST/v2/auditsSubmit a batch for scoringyes*
    GET/v2/audits/{id}Fetch batch status and summaryyes
    GET/v2/audits/{id}/resultsStream per-conversation results (NDJSON)yes
    POST/v2/rubricsRegister a rubric versionyes*
    GET/v2/rubrics/{id}Fetch a rubric definitionyes
    POST/v2/rubrics/{id}/versionsBump a rubric's versionno

    * Idempotent with an Idempotency-Key request header; see §2.3.

    2.2 POST /v2/audits

    Submit a batch of conversations against a rubric. Returns a batch handle; scoring proceeds asynchronously.

    rubric_id
    Rubric UUID. Must reference a rubric version the caller has read access to.
    rubric_version
    Optional. Defaults to latest. Pinned versions are recommended for reproducibility.
    conversations[]
    Array of {id, transcript}. Transcripts are a linear list of {role, text, timestamp} turns.
    callback_url
    Optional HTTPS URL; receives a batch.completed POST when scoring finishes.
    Request · POST /v2/auditshttp
    POST /v2/audits HTTP/1.1
    Host: api.ontopix.ai
    Content-Type: application/json
    Idempotency-Key: 01JC8YW2R3PQ1N1NQ6CV1Z0KFP
    
    {
      "rubric_id": "7c5b1d4e-…",
      "rubric_version": "3.2.0",
      "conversations": [
        { "id": "c_001", "transcript": […] },
        { "id": "c_002", "transcript": […] }
      ],
      "callback_url": "https://customer.example/audibot/hook"
    }
    Response · 202 Acceptedjson
    {
      "batch_id": "b_01JC8YWR9K…",
      "status": "queued",
      "conversation_count": 2,
      "rubric_version": "3.2.0",
      "created_at": "2026-04-23T09:12:04Z"
    }
    ✓ Tip

    Pin rubric_version in production. Leaving it at latest silently changes scoring behavior the next time the rubric is republished, which defeats trend analysis across consecutive batches.

    SPEC-AUDIBOT-v2 · 2.0.0-draft.3 Ontopix · Confidential
    Audibot v2 API · Specification 3. Data model 5

    3. Data model

    3.1 Rubric

    A rubric is a versioned, immutable list of criteria. Each criterion has a weight, a prompt the evaluator LLM uses to score, and a scale. Weights within a rubric must sum to 1.0 ± 0.001.

    Rubric fields.
    FieldTypeNotes
    iduuidAssigned by server on first POST.
    versionsemverImmutable per (id, version) pair.
    criteria[]Criterion[]Ordered; display order preserved.
    min_score, max_scorefloatScale bounds for the aggregate score.
    created_atdatetimeRFC 3339 UTC.
    Criterion schematypescript
    interface Criterion {
      // Stable across versions of the rubric.
      id: string;
      name: string;
      weight: number;      // 0..1; rubric weights sum to 1
      prompt: string;     // LLM instruction
      scale: {
        type: "ordinal" | "continuous";
        labels?: string[];  // required for ordinal
        range?: [number, number]; // required for continuous
      };
    }

    3.2 AuditResult & Evidence

    Each AuditResult carries evidence per criterion: one or more transcript spans the evaluator relied on, plus a short reasoning string. Disputes are resolved by pointing at the evidence, not by re-running the scorer.

    AuditResulttypescript
    interface AuditResult {
      conversation_id: string;
      rubric_version: string;
      overall_score: number;
      criterion_scores: {
        criterion_id: string;
        score: number;
        evidence: Evidence[];
      }[];
    }
    
    interface Evidence {
      turn_index_start: number;
      turn_index_end: number;
      excerpt: string;        // verbatim text
      reasoning: string;      // ≤ 240 chars
    }
    ℹ Note

    Evidence excerpts are verbatim substrings of the turn text at turn_index_start..turn_index_end. Consumers can diff them against their own transcript store to verify that no span has drifted between submission and result delivery.

    SPEC-AUDIBOT-v2 · 2.0.0-draft.3 Ontopix · Confidential
    Audibot v2 API · Specification 4. Versioning & migration 6

    4. Versioning & migration from v1

    v1 and v2 coexist at /v1/* and /v2/* respectively through 2026-09. After that date v1 returns 410 Gone with a Sunset header pointing at this spec2.

    4.1 Breaking changes

    Key breaking deltas v1 → v2.
    Areav1v2
    Rubric shapeFlat scoresheet (single score per conversation)Weighted criteria, per-criterion scores and evidence
    Scoring modeSync (POST /evaluate) + async batchesAsync batches only
    EvidenceFree-text blobStructured transcript spans + verbatim excerpt
    Rubric versioningImplicit; latest appliesExplicit semver; pinning recommended
    Result deliverySingle JSON envelopeNDJSON stream from /audits/{id}/results

    4.2 Migration path

    Recommended order of operations for integrators migrating in place:

    1. Re-author rubrics against the v2 schema. Preserve criterion IDs from v1 where the semantics are unchanged — v2 honors the ID as a stable external reference.
    2. Pin rubric_version in all production calls.
    3. Switch to async batches. Replace sync-only code paths with the submit + webhook pattern.
    4. Update result ingestion to NDJSON streaming; v2 results may arrive out of order.
    5. Decommission v1 rubric registrations before 2026-09.
    ◇ Rationale

    Preserving criterion IDs across the migration is explicit rather than automatic because we cannot verify that the meaning of a criterion is preserved across a rubric rewrite — only the caller can. Making it the caller's choice keeps the audit trail clean: a v2 criterion with the same ID as a v1 criterion is an explicit claim of semantic continuity.

    The full comparison matrix and per-endpoint diff are published alongside this spec at packages/channels/document/examples/ in the stylebook repository.

    1. Scoring-model weights: see ADR-017 · Audibot model pinning. Not part of the v2 contract.
    2. See RFC 8594 for the Sunset header semantics.
    SPEC-AUDIBOT-v2 · 2.0.0-draft.3 Ontopix · Confidential