# Langfuse Agent Tracing: Tools, Retries, and Failures

> A tested Python pattern for Langfuse agent traces with explicit tool parentage, separate retries, child errors, redaction, and exporter isolation.
> Author: Roman Belov · Published: 2026-09-26 · Source: https://futurecraft.pro/blog/langfuse-agent-tool-tracing/

Our first remote Langfuse trace contained the root task and none of its children. The
workflow had completed correctly. Six adapter tests were green. The fault was one
attribute name: the installed Python SDK wrapper exposed the span ID as `id`, while the
contract we had copied from a class docstring used `observation_id`.

That failure is why this is not another diagram assembled from documentation. We ran a
fixed support workflow against a self-hosted Langfuse instance, inspected the returned
parentage, and repeated it for a successful run, a tool retry, and a terminal tool
error. We also sent the same workflow to an unreachable exporter.

The pattern that survived those tests is straightforward:

> Create one root observation for the user-visible task. Record retrieval, model work,
> and each tool attempt as a child. Keep tool failure on the tool observation, summarize
> the task outcome on the root, and redact before telemetry leaves the application.

This article stays on that implementation problem. For installation and credentials,
use the [current Langfuse setup tutorial](/blog/langfuse-step-by-step/). For metrics,
alerts, and the larger operating model, start with
[production LLM observability](/blog/llm-observability-langfuse/).

## What we tested

The fixture answers a synthetic refund question. It retrieves one policy, records a
plan, calls two deterministic tools, and composes an answer:

```text
answer-support-question
├── retrieve-account-policy
├── draft-plan
├── tool:get-subscription
├── tool:get-refund-window
└── compose-answer
```

The run used no live model or customer data. Model responses and tool results were
fixed so that a rerun would test trace structure rather than provider drift.

| Component | Version in the run |
| --- | --- |
| Langfuse Python SDK | `4.14.5` |
| Self-hosted Langfuse | `3.205.0` |
| Python | `3.13.9` |
| Fixture release | `6401879` |

We ran three remote scenarios on 26 August 2026:

1. both tools succeed;
2. the subscription tool times out, then succeeds on its second attempt;
3. the refund tool ends with a terminal error and the answer is sent for human review.

A fourth process pointed Langfuse at `127.0.0.1:9`. The application still returned the
same fixture answer. That test did not prove that telemetry is never lost; it proved
the narrower claim that this workflow did not make telemetry availability part of its
result path.

![Sanitized Langfuse trace shape showing separate retry attempts](/artifacts/langfuse-agent-tool-tracing/sanitized-trace-shape-2026-08-26.svg)

## Give the task one root observation

The root should represent the transaction a user would recognize: answer a support
question, prepare an itinerary, or review a pull request. If the root represents only
one model call, tool failures and recovery happen outside the task boundary.

Langfuse documents a trace as the observations that share one trace ID. It also maps
observations to OpenTelemetry spans and supports nested context propagation. The
[SDK overview](https://langfuse.com/docs/observability/sdk/overview) covers that model.
The current
[Python instrumentation guide](https://langfuse.com/docs/observability/sdk/instrumentation)
documents explicit `trace_context` with `trace_id` and `parent_span_id`.

Automatic context is convenient inside one synchronous call stack. Our fixture uses
explicit parentage because the same adapter boundary can later receive events from a
queue or another service. The important part is the SDK wrapper's actual `trace_id`
and `id` values:

```python
parent = self.observations[event.parent_id]
trace_context = {
    "trace_id": parent.trace_id,
    "parent_span_id": parent.id,
}

observation = self.client.start_observation(
    name=event.name,
    as_type=event.kind,
    trace_context=trace_context,
    metadata=safe_metadata(event),
    version=self.fixture_version,
    model="recorded-fixture-v1" if event.kind == "generation" else None,
)
```

The first implementation read `parent.observation_id`. Our fake object exposed that
same invented field, so the contract tests confirmed each other instead of confirming
the installed SDK. In the remote run, child creation failed inside telemetry isolation
and only the root arrived.

The fix was small: use `parent.id`. The lesson is broader but not complicated. If an
adapter controls causal parentage, run at least one integration test against the exact
SDK version you ship. A fake should copy the runtime surface, not the documentation you
hope is accurate.

## Record every retry as a separate tool attempt

Do not reopen or overwrite the failed tool observation. A retry is another network
attempt, with its own start, end, error, and result.

Our timeout scenario produced this shape:

```text
answer-support-question                 DEFAULT
├── tool:get-subscription  attempt=1    ERROR   TimeoutError
├── tool:get-subscription  attempt=2    DEFAULT
├── tool:get-refund-window attempt=1    DEFAULT
└── other successful children           DEFAULT
```

Both subscription attempts share the root parent but have different observation IDs.
The first ends before the second begins. That gives an operator enough evidence to
answer three practical questions: did a retry happen, which attempt failed, and did the
task recover?

One observation with `attempt=2` and a successful final status would hide the timeout.
One observation left in `ERROR` after recovery would make the task look worse than the
user experienced. Separate attempts preserve both facts.

In application code, generate a logical ID per attempt rather than per tool name:

```python
for attempt in (1, 2):
    observation_id = f"subscription-{attempt}"
    # Application wrapper, not a Langfuse SDK method.
    trace.start_tool(
        observation_id=observation_id,
        name="tool:get-subscription",
        parent_id=root_id,
        attempt=attempt,
    )
    try:
        subscription = tools.get_subscription(account_id)
    except TimeoutError as error:
        trace.end_tool(
            observation_id=observation_id,
            level="ERROR",
            error_type=type(error).__name__,
        )
        if attempt == 2:
            raise
        continue

    trace.end_tool(observation_id=observation_id, level="DEFAULT")
    break
```

The downloadable fixture uses a small sink interface rather than the illustrative
`start_tool` wrapper above. This keeps the agent workflow independent of Langfuse and
makes exporter failure testable.

## Keep child failure and task outcome separate

A failed tool call does not always mean the whole task failed. It may be retried, or
the application may return a safe fallback that needs human review.

In the terminal scenario, `tool:get-refund-window` ended as `ERROR`. Retrieval, the
subscription lookup, and both generations remained `DEFAULT`. The root ended as
`WARNING`, because the workflow returned a bounded answer but could not confirm refund
eligibility.

```text
tool:get-refund-window → ERROR: TerminalToolError
answer-support-question → WARNING: requires human review
successful siblings → DEFAULT
```

This is an engineering policy, not a Langfuse law. Define the mapping before adding
dashboards:

| Observation | Level | Meaning in this fixture |
| --- | --- | --- |
| Failed tool attempt | `ERROR` | That operation did not complete |
| Recovered root | `DEFAULT` | The task produced its intended result |
| Degraded root | `WARNING` | The task returned a safe but incomplete result |
| Successful sibling | `DEFAULT` | Its own operation completed |

Without that distinction, an alert on child errors can be mistaken for a count of
failed user tasks. The broader
[production LLM stack control plane](/blog/production-llm-stack/) should track both
operation-level failure and task-level outcome.

## Redact before the SDK sees the event

The fixture starts with a synthetic email address and account ID. Its redaction
function runs while trace attributes are created, before the adapter calls Langfuse.
No observation input or output payload is exported at all.

```python
EMAIL_PATTERN = re.compile(
    r"[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}",
    re.I,
)
ACCOUNT_PATTERN = re.compile(r"\bacct_[A-Z0-9]+\b", re.I)


def redact_text(value: str) -> str:
    value = EMAIL_PATTERN.sub("[EMAIL_REDACTED]", value)
    return ACCOUNT_PATTERN.sub("[ACCOUNT_REDACTED]", value)
```

These two expressions cover the synthetic fixture, not every form of personal data.
Production redaction needs a data inventory and tests for the identifiers your product
actually handles. Langfuse's current
[masking documentation](https://langfuse.com/docs/observability/features/masking)
recommends masking before trace data leaves the application and documents SDK hooks for
inputs, outputs, metadata, and OpenTelemetry attributes.

We added a second boundary for the published evidence. The authenticated trace response
contained the Langfuse public key inside SDK scope metadata. A public key is not the
secret key, but it still did not belong in a portable trace sample. The exporter now
copies only a fixed set of fields and rejects either credential if it appears in the
saved JSON.

Redaction and export minimization solve different problems:

- redaction changes sensitive values before telemetry is sent;
- an allowlist prevents unrelated server metadata from entering a public artifact;
- a final negative scan checks the exact file that will be published.

## Keep telemetry outside the result path

The workflow accepts a `TraceSink`. A `SafeTraceSink` catches exporter-specific
exceptions and counts dropped events without changing the business result:

```python
@dataclass
class SafeTraceSink:
    delegate: TraceSink
    dropped_events: int = 0

    def record(self, event: TraceEvent) -> None:
        try:
            self.delegate.record(event)
        except Exception:
            self.dropped_events += 1
```

Langfuse states that SDK errors are caught and logged, and that requests are exported
asynchronously. Its
[Python instrumentation guide](https://langfuse.com/docs/observability/sdk/instrumentation)
also tells short-lived processes to call `flush()`. Our application-level guard does
not replace those SDK guarantees; it keeps the workflow contract independent of a
specific exporter.

There is a cost to this choice: traces can be incomplete while the user still receives
an answer. Track dropped telemetry through a separate operational signal. Do not turn
an observability outage into a product outage merely to make a dashboard complete.

## Reproduce the evidence

The [artifact bundle](/artifacts/langfuse-agent-tool-tracing/) contains the deterministic
workflow, Langfuse adapter, standard-library tests, integration runner, sanitized JSON,
and the diagram shown above. Code is MIT-licensed; the evidence report and diagram use
CC BY 4.0.

Run the dependency-free contract suite first:

```bash
# Run this from the directory containing the downloaded files.
python3 -m unittest discover -s . -p 'test_*.py' -v
```

For a real integration run, provide dedicated project credentials through environment
variables and use an isolated SDK environment:

```bash
uv run --isolated --with langfuse==4.14.5 python \
  run_integration.py \
  --output /tmp/futurecraft-langfuse-evidence.json
```

Do not print credentials or save the raw authenticated response. The runner minimizes
the response in memory, checks parentage and levels, scans for synthetic identifiers and
keys, and writes only the sanitized shape.

## What this run does not prove

- The recorded model and tools do not represent live provider behavior.
- Millisecond timestamps from one local run are not a latency benchmark.
- One SDK and server combination does not prove compatibility with every release.
- The fixture's regexes are not a universal PII detector.
- A complete trace does not prove that the model answer is correct.

The useful result is narrower: with the tested versions, explicit parentage produced
one inspectable task tree; retries remained separate; a terminal child error did not
erase successful work; redaction happened before export; and an unavailable exporter
did not change the fixture result.

That is enough to turn agent tracing from a decorative waterfall into an operational
record you can challenge and rerun.
