OpenTelemetry-native tracing for AI agents and LLM applications. glassflow-rius
emits OpenTelemetry GenAI
traces over OTLP to the managed GlassFlow observability platform (or any
OTLP-compatible backend).
Status: alpha. APIs may change.
pip install glassflow-riusimport rius
from rius import observe, start_as_current_generation, start_as_current_span
from rius.semconv import SpanKind
rius.init(
api_key="glassflow_...", # or set RIUS_API_KEY
service_name="my-agent", # or set RIUS_SERVICE_NAME
)
# 1. Decorator — trace a whole function
@observe
def handle(query: str) -> str: ...
# 2. Context manager — trace a block
with start_as_current_span("retrieve", kind=SpanKind.RETRIEVER) as obs:
obs.set_output(docs)
# 3. LLM generations — gen_ai-native
with start_as_current_generation("chat", model="gpt-4o", input=messages) as gen:
gen.set_output(reply)
gen.set_usage(input_tokens=42, output_tokens=17)Span names are optional. Left out, a span is named the way the OpenTelemetry
GenAI conventions say it should be — the operation followed by what it acted on:
chat gpt-4o, embeddings text-embedding-3-small, execute_tool get_weather,
invoke_agent planner, retrieval product-kb — degrading to the bare operation
(chat) when that identifier is unknown. @observe keeps the function's
qualified name for a plain CHAIN step, which has no operation; an unnamed
start_span() of that kind is called chain. An explicit name always wins.
Each surface has a manual variant for lifetimes a with block can't express
(streaming, callbacks): start_span(...) / start_generation(...) return a handle
you .update() and must .end() yourself. A manual handle does not auto-record
exceptions, so report a failure with .record_exception(exc), which records the
exception event, sets the ERROR status and sets error.type together:
gen = start_generation("chat", model="gpt-4o", input=messages)
try:
reply = call_the_model(messages)
except Exception as exc:
gen.record_exception(exc)
raise
finally:
gen.end()Configuration is resolved from explicit arguments first, then environment variables:
| Argument | Environment variable | Default | Description |
|---|---|---|---|
endpoint |
RIUS_ENDPOINT |
https://ingest.eu.console.rius-glassflow.com |
Base OTLP endpoint. Traces are sent to <endpoint>/v1/traces. |
api_key |
RIUS_API_KEY |
— | Injected as an Authorization: Bearer <key> header on every export. |
service_name |
RIUS_SERVICE_NAME |
unknown_service |
Sets the OpenTelemetry service.name resource attribute. |
disabled |
RIUS_DISABLED |
false |
Kill switch. When true, spans are created but never exported. |
sample_rate |
RIUS_SAMPLE_RATE |
1.0 |
Head sampling ratio 0.0–1.0 (whole-trace; children follow root). |
capture_content |
RIUS_CAPTURE_CONTENT |
true |
When false, prompt/response content is stripped at export (metadata still sent). |
bridge_foreign_provider |
RIUS_BRIDGE_FOREIGN_PROVIDER |
false |
Also export the spans of another OpenTelemetry SDK that owns the global provider (see below). |
mask is a code-only option (no env var): pass a callable to init(mask=...) and
it is applied to every content attribute value at export, across our spans and any
bundled third-party instrumentation. A mask that accepts a key keyword also
receives the attribute key, for per-attribute decisions.
Both controls cover the whole span, not just its attributes: event and link
attributes, the exception.message / exception.stacktrace of a recorded
exception, and the ERROR status description, since provider errors routinely
echo the rejected request back. With capture_content=False those are
stripped and the status code and exception.type are kept, so failures stay
visible; with a mask, the mask runs over them (the status description is
offered under key="status.description").
rius.init(mask=lambda value: "[REDACTED]") # redact all captured content
rius.init(mask=lambda value, *, key: hash_pii(value) if "input" in key else value)
rius.init(capture_content=False) # drop content entirely, keep metadataOpenTelemetry has one global tracer provider per process. If another SDK
(Langfuse v3, for one) claims it before rius.init(), your own job and tool
spans go to that SDK only, while the LLM spans inside them still reach Rius, so
Rius receives traces whose parents it never gets. rius detects this: those LLM
spans carry rius.parent.foreign=true, the resource records
rius.sdk.global_provider=foreign:<class>, and the first one logs a warning.
To send the other SDK's spans to Rius as well, opt in with
init(bridge_foreign_provider=True) or RIUS_BRIDGE_FOREIGN_PROVIDER=true.
rius then exports that SDK's spans too, subject to rius's sample_rate,
capture_content and mask, and it never modifies what the other SDK
exports: the other vendor receives exactly what it would without rius.
The SDK bundles existing OTel instrumentors (OpenInference) as optional extras,
so a single install captures your LLM provider and framework calls. Install the
extras you need and init() enables whatever it finds; the instrumentation
spans nest under your @observe / start_as_current_span traces automatically.
pip install "glassflow-rius[openai]" # one provider
pip install "glassflow-rius[instruments]" # everything supportedAn extra installs the instrumentation for a library, never the library
itself, so the SDK never pins or upgrades the versions your code runs against:
glassflow-rius[anthropic] expects anthropic to be your own dependency,
which it already is in any project that calls Anthropic. An extra installed
without its library makes init() log a DependencyConflict from
OpenTelemetry and leaves that one integration off.
rius.init() # auto-enables installed instrumentors
rius.init(instruments=["openai"]) # restrict to specific ones
rius.init(instruments=[]) # disable auto-instrumentationSupported instruments: openai, anthropic, langchain, llama-index,
litellm, and mcp. Content captured by instrumentors is covered by the same
mask / capture_content controls as our own spans.
The mcp instrument is built in and has no extra: there is no third-party
instrumentation to install, and an extra that installed mcp for you would
break the rule above. If the mcp
package is present, every ClientSession.call_tool() your agent makes becomes
a first-class TOOL span (execute_tool <name>) with gen_ai.tool.name, the
arguments and result as input.value/output.value, latency, and error status
(including tools that return isError results).
Instrumentors patch libraries process-wide, so a scoped client
(init(set_global=False)) only enables them when instruments=[...] is passed
explicitly. Calling init() again while a client is active logs a warning and
returns the existing client unchanged; call client.shutdown() first to
reconfigure.
OpenTelemetry caps a span at 128 attributes by default and, once a span is full, silently evicts its OLDEST keys. The OpenInference instrumentors write one attribute per message field and per tool field, so an agent loop with ten tools passes 128 within about six (Anthropic) to nine (OpenAI) turns, and what goes first is what they wrote first: the model and request parameters, the tool definitions, and the system prompt.
init() therefore gives its tracer provider a limit of 4096 attributes per
span. The value-length limit is left alone (the SDK caps its own JSON
attributes). If you set OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT or
OTEL_ATTRIBUTE_COUNT_LIMIT to a non-negative integer, your value wins (a blank value
is ignored, and an invalid one is ignored with a warning) and OpenTelemetry resolves the
limits as usual. A span that still hits the limit reports how many attributes
it dropped (dropped_attributes_count on the wire), so a loss is visible
rather than silent.
The limit applies to the provider init() builds. init() does not accept a
provider of your own; if you run OpenInference instrumentors without this
SDK, on your own TracerProvider, raise the limit yourself, either with
OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT=4096 in the environment or in code:
from opentelemetry.sdk.trace import SpanLimits, TracerProvider
provider = TracerProvider(span_limits=SpanLimits(max_span_attributes=4096))Export is designed to never block or crash your application:
- Async batched export. Spans are queued in-process and exported in batches
from a background thread (
BatchSpanProcessor). Span creation stays fast even when the backend is slow or unreachable. - Retries. Transient failures (connection errors, 429/5xx) are retried with exponential backoff and jitter, bounded by the export timeout.
- Graceful degradation. If the backend stays down, spans are dropped and an
error is logged — exceptions never propagate into application code. A failing
maskcallable drops only the affected attribute value (fail closed), never the batch. - Flush on shutdown. Pending spans are flushed automatically at interpreter
exit. Call
client.flush()to force an export, orclient.shutdown()to drain and stop.
Batching and backpressure are tunable via the standard OpenTelemetry env vars:
OTEL_BSP_MAX_QUEUE_SIZE (default 2048; spans beyond this are dropped),
OTEL_BSP_SCHEDULE_DELAY (default 5000 ms), OTEL_BSP_MAX_EXPORT_BATCH_SIZE
(default 512), and OTEL_BSP_EXPORT_TIMEOUT (default 30000 ms).
uv sync --group dev
uv run pytest
uv run ruff check . && uv run ruff format --check .
uv run mypyReleases are automated with release-please and published to PyPI via Trusted Publishing.
- Merge changes to
mainusing Conventional Commits (feat:→ minor,fix:→ patch,feat!:/BREAKING CHANGE→ major). - release-please keeps a Release PR open that bumps
__version__and updatesCHANGELOG.md. Merge it when you want to cut a release. - Merging tags
vX.Y.Z, creates a GitHub Release, and publishes to PyPI automatically.
Non-conventional commits are ignored for versioning.
MIT