Listening to Events
Both inference and embeddings runtimes expose two ways to listen to events:Targeted Listeners
UseonEvent() to listen for a specific event class:
Wiretap
Usewiretap() to receive all events regardless of type. This is useful for debugging and general-purpose logging:
Inference Events
The inference lifecycle dispatches events in this order:Execution-Level Events
These events bracket the entire inference operation, including any retry attempts.
InferenceCompleted is dispatched exactly once per execution, whether it succeeded or failed.
Attempt-Level Events
Each retry attempt dispatches its own events:
When retries are configured, you may see multiple
InferenceAttemptStarted/InferenceAttemptFailed pairs before a final InferenceAttemptSucceeded event. The attemptNumber field tracks which attempt is running.
Response Events
InferenceResponseCreated is emitted from two places — the driver, for a non-streamed response, and InferenceStream, when a stream finalises. The key set is the same either way; both use a single payload builder (Inference\Core\InferenceResponseEventPayload), so the two cannot drift.
One value does differ. data['executionId'] is populated on the streamed path and is null on the non-streamed path — the driver is shared across executions by InferenceRuntime, so it has no execution to name. Correlate a non-streamed response by data['requestId'], which both paths carry and which InferenceStarted and InferenceCompleted report alongside their executionId.
data['statusCode'] is omitted, on both paths alike, when the response carries no HTTP status — for example a response assembled purely from stream deltas.
Streaming Events
The
StreamFirstChunkReceived event is particularly useful for measuring time-to-first-chunk (TTFC), as it includes the requestStartedAt timestamp.
Driver Events
Sensitive configuration values (API keys, tokens, secrets) are automatically redacted in the
InferenceDriverBuilt event payload.
Embeddings Events
The embeddings lifecycle dispatches a smaller set of events:Practical Examples
Logging Token Usage
Measuring Time-to-First-Chunk
Tracking Retry Attempts
Monitoring Execution Outcomes
Event Dispatcher
Events are dispatched through anEventDispatcher that implements CanHandleEvents (which extends Psr\EventDispatcher\EventDispatcherInterface). When a runtime is created without an explicit event dispatcher, it creates a default one named 'polyglot.inference.runtime' or 'polyglot.embeddings.runtime'.
You can inject a shared event dispatcher to correlate events across multiple runtimes or integrate with your application’s existing event system:
Listener Gating
Some events are not constructed at all when nothing is listening for them. Building an event is not free — the baseEvent generates a UUID and a DateTimeImmutable (~0.9µs), and the payload arrays cost more on top: InferenceRequested walks the message list, tools and options; InferenceResponseCreated runs strlen() over the full content; the InferenceFailed payloads run header/body redaction. On the streaming path the per-delta events would pay that cost thousands of times per response.
Gating is decided by Cognesy\Events\Support\ListenerGate, the single definition of the rule:
Fail-open is contractual
A dispatcher is only asked about its listeners if it implementsCognesy\Events\Contracts\CanCheckListeners. A plain PSR-14 EventDispatcherInterface cannot report its listeners, so it is assumed to listen and receives every event. No dispatcher ever loses an event to this optimisation — the worst case is that the payload is built and discarded. The built-in EventDispatcher does implement CanCheckListeners, so it gets the gating.
The gate is resolved once, at construction
Emitters resolve the answer in their constructor and store it in areadonly bool; they do not re-check per dispatch, because that would put an instanceof back on the hot path. The deliberate consequence:
A listener registered after the emitter was constructed is not observed by that emitter.For a
wiretap() or onEvent() call made on the runtime before the request is sent, this is invisible — the emitters do not exist yet. It becomes visible if you register a listener from inside another listener mid-request, or attach one to a shared dispatcher while a stream is already being consumed: gated events already resolved as “unwanted” stay unwanted for the rest of that stream. Register listeners before starting the operation you want to observe.
What is gated
Everything else is dispatched unconditionally.
packages/instructor applies the same rule to
its own structured-output lifecycle emitters — see the “Listener Gating” section of
packages/instructor/docs/internals/events.md.
The lifecycle events matter most. Each one carries a telemetry envelope under
data['telemetry'], and four of the six build it by serialising the entire conversation
via Messages::toArray(). That cost scales with conversation length, not response length:
for a 128 KB conversation it was ~534 µs per request. With no listeners a session now costs
a flat ~14 µs regardless of how long the conversation is.
When listeners are attached, the conversation is serialised once per request rather than
four times. Cognesy\Polyglot\Telemetry\MessagesSerializationMemo holds the two most recent
results in fixed slots keyed on the Messages instance, so the four envelope sites share
one serialisation — ~166 µs instead of ~534 µs under a wiretap on a 128 KB conversation. The
key is the object, not its content: a conversation rewritten mid-session by length recovery
is a different instance and is serialised afresh.
Two slots rather than one because packages/instructor uses the same memo and interleaves
two conversations per request — its own, and the materialized one handed to the nested
inference call. Across a structured-output request that takes the count from 7 to 2. Deeper
nesting simply misses and re-serialises: a lost optimisation, never a wrong answer.
Only the payload and the dispatch are conditional. Timing, attempt numbering and execution
state are not — durationMs is correct whether or not anyone was listening.