Every APM agent ever written assumes your process outlives the request. It starts a background thread, buffers spans in memory, and flushes on a timer. That assumption does all the work, and serverless deletes it.
In Lambda, the runtime freezes the execution environment the moment your handler returns. Not exits — freezes. Your flush thread stops mid-instruction, and whatever sits in the buffer stays there until the environment thaws, which might be in eight minutes, or never, because the environment gets reclaimed. You do not lose a percentage of your telemetry. You lose the tail of every invocation, in the pattern that best hides low-traffic errors: the functions invoked least are reclaimed before they flush.
Edge runtimes are stricter. On Cloudflare Workers or Vercel Edge there is no filesystem to install an agent onto, no native binary, no require hook to patch, and a bound on CPU after the response is sent. So the serverless question is not which vendor has the best dashboard. It is how the telemetry physically gets out.
Key takeaways
- No persistent process means no background flush. Telemetry is either flushed synchronously before the handler returns, adding latency to every request, or handed to a separate process that survives the freeze.
- Lambda extensions are the standard answer: a co-process receiving the Telemetry API stream. They work, and they add cold start time.
- Cold start attribution is the metric everyone wants and few tools separate cleanly. Demand init duration split from handler duration.
- Queue-triggered invocations lose parent context unless it travels in the message. Batch consumption needs span links, not a single parent.
Why the agent model breaks
Three properties of the execution model each break a different part of a conventional agent.
The process is frozen, not idle, between invocations. Anything on setInterval, a background goroutine, or a Java daemon thread does not run. Code written to flush “every five seconds” flushes when the next invocation thaws the environment, by which point the data may be minutes stale or gone.
The environment is disposable and unnamed. Agents that identify by hostname have nothing to report. There is no host, which is also why per-host pricing makes no sense here.
Init time is billed and visible. An agent taking 800ms to start is a deployment concern you notice once on a long-lived service. In Lambda that cost recurs on every cold start and lands in your p99.
Three ways out. Flush in-request with forceFlush() before returning: correct, simple, and you pay an export round trip inside billed duration — for a 40ms function that is not a rounding error. Ship out of band via a co-process, which is what Lambda extensions do. Ship via logs, writing structured JSON to stdout: no handler latency, but you pay log ingestion prices and reconstruct spans from lines.
Needs first-hand data: Deploy the same function three ways — synchronous flush, extension export, log-based export — and measure billed duration at p50 and p99 for each, plus cold start init duration. The synchronous penalty and the extension’s init cost decide this, and both are workload-specific.
Lambda layers and extensions, and what they cost
Vendors ship serverless support as a Lambda layer: a zip containing an extension binary and often the SDK, added by ARN, plus environment variables and sometimes an AWS_LAMBDA_EXEC_WRAPPER script that sets up the runtime before your handler loads.
The layer is extracted into /opt. The extension registers with the Extensions API during init, subscribes to the Telemetry API, and opens a local HTTP listener. Your handler exports OTLP to localhost — a loopback write, no network latency — and the extension batches, forwards, and flushes on shutdown.
Right architecture, and my default. Be clear about the cost. The extension initializes during init, so it is on the cold start path. It ships with your deployment package. It consumes memory from the function’s allocation, a meaningful fraction at 128MB. Auto-instrumentation wrappers add their own init work — the same startup penalty Java APM agents impose, except you pay it repeatedly. And the shutdown window is bounded: if the export endpoint is slow, that data is dropped rather than retried.
Cold start attribution is the metric everyone asks for
Every serverless team wants one number: how much of my latency is cold starts. That is harder to produce than it sounds.
A cold start has phases that are not equally your problem. Package download and extraction scales with package size. Runtime bootstrap is the language runtime starting — why a JVM function and a Go function are in different worlds. Init code is everything at module scope: SDK clients, pools, config, and your APM layer. Then the handler runs. Only the last two are yours to fix, and they must be separated to be actionable.
What to demand from a tool:
- Init duration separate from handler duration, not summed into one “cold start latency” number.
- Cold start rate by function version and memory size — memory determines CPU, so a slow-to-init function is often fixed with more memory, not less code.
- Cold start flagged on individual traces, so a p99 outlier is classified at a glance.
- Provisioned concurrency utilization, since paying for warm environments you do not use is invisible waste.
The trap: a tool reporting average duration across warm and cold invocations tells you nothing, because the two distributions are so different that the mean sits where no invocation lives.
Needs first-hand data: For your three highest-volume functions, record cold start rate and init duration at two memory settings each. Memory buys CPU, so init-heavy functions often get cheaper as you raise it — the duration drop can outweigh the per-ms price increase.
Async invocations are where traces quietly break
Synchronous tracing is solved. API Gateway carries traceparent, the handler extracts it, spans nest correctly. Asynchronous invocation is where it falls apart, and the failure is silent — complete-looking traces missing everything downstream.
SQS. Trace context travels in message attributes. A producer that does not add it, or a consumer that does not extract it, produces two unrelated traces. Worse, batch consumption gives one invocation N messages with N parent contexts, so the correct model is a span with multiple links, not a single parent. Many tools flatten this and pick the first message’s context.
SNS and EventBridge. Same problem, different attribute mechanism, and any payload transformation in between can drop the attributes you relied on.
Step Functions. State transitions are orchestrated by the service, not your code, so without integration into execution history you get disconnected traces and no view of the state machine. Direct async invoke queues internally and returns immediately, so the parent span ends before the child starts.
The fix is span links rather than parent-child for anything queue-mediated, plus explicit propagation in the message. When evaluating, send five messages through SQS, consume them in one invocation, and look at the trace. That test separates the tools that understand async from the ones that assume HTTP.
Edge runtimes: no agent, fetch only
Cloudflare Workers, Deno Deploy, and Vercel Edge are not small Lambdas. They are V8 isolates rather than containers, with no filesystem, no native modules, and no process to attach to.
- No binary agent, ever. All instrumentation is library code inside the isolate.
- Export is HTTP only. OTLP over
fetch. No gRPC, no UDP, no local socket. - CPU after the response is limited. Cloudflare’s
waitUntilextends work past the response, which is where export belongs, but the runtime bounds it. A slow collector means dropped telemetry. - Instrumentation is manual. No
requirehook to patch, so Node-style auto-instrumentation barely applies. You wrap your ownfetchcalls. - Cardinality is a live risk. A colo label on every metric multiplies series count by the size of the network — the same trap as Kubernetes with a different label.
Practically, edge observability today means the OpenTelemetry JS SDK with an OTLP-over-HTTP exporter invoked from waitUntil. That constraint is the single most useful filter on this list: a vendor with an excellent Lambda layer and no OTLP-over-HTTP endpoint is unusable at the edge, and Lambda support tells you nothing about Workers support.
Log-based versus trace-based
Log-based means structured JSON to stdout carried by the platform’s log pipeline. Zero handler latency, no layer, identical across every runtime including edge — sometimes the only option. You pay log ingestion prices for what is really trace data and reconstruct causality by joining on IDs.
Trace-based means real spans with parent-child structure, exported via an extension or synchronous flush. Proper waterfalls, service maps, latency breakdowns, paid for in cold start time, package size, and either handler latency or extension memory.
My rule: trace-based for the synchronous request path, log-based for fire-and-forget, low-volume, or anywhere you cannot install a layer. Emitting a log line with the trace ID from every function is cheap and lets you jump between views.
OpenTelemetry: the standard underneath, not a product

OpenTelemetry is not one of the products below — it is the instrumentation standard most of them consume. It is a set of SDKs, semantic conventions, and a Collector. There is no dashboard, no alerting, no vendor, and no bill, so ranking it against Datadog or Lumigo is a category error. What it does decide is portability: building on OTel constrains which products you can later leave, because the instrumentation stops being the vendor’s property.
OpenTelemetry publishes Lambda layers for the major runtimes, giving you the same extension pattern with vendor-neutral OTLP export: an extension process registers with the Extensions API, your handler writes OTLP to loopback, and the collector inside the layer forwards on. It is also the only viable path at the edge, where the JS SDK with an OTLP-over-HTTP exporter called from waitUntil is essentially the entire available toolkit.
What it gives you
- Same layer-and-extension architecture as the commercial options with no vendor coupling — the export destination is one environment variable
- The only realistic instrumentation for V8 isolate edge runtimes, where no binary agent can exist
- The collector inside the layer is configurable, so you can batch, filter and redact before anything leaves the function
- No per-invocation licensing, so instrumenting low-value high-frequency functions costs nothing beyond the telemetry itself
What it does not do
- The bundled collector adds init work on the cold start path and consumes function memory, and tuning that down is your job
- No AWS service-level understanding out of the box — Step Functions execution history and payload reconstruction are not there
- SQS batch handling with span links depends on the instrumentation library version and the backend rendering them properly
- Stores nothing and shows nothing: you still need a backend to receive the OTLP stream, and the cold start milliseconds and function memory the collector consumes are cost on top of it
Datadog

Datadog has the most complete Lambda story in the category. Its extension also handles log forwarding and metric submission, so you run one shipping path rather than three — no separate log subscription filter, no separate metric agent, one layer ARN and a set of environment variables. It reads the Telemetry API for init and shutdown phases, which is what makes its cold start reporting usable rather than approximate.
Pros
- One extension covers traces, logs and custom metrics, so you are not maintaining three export paths per function
- Init duration reported separately from handler duration, which is the split that makes cold start data actionable
- Cold starts flagged on individual traces, so a p99 outlier is classified without manual correlation
- Broad coverage of the AWS services around the function — SQS, Step Functions, API Gateway — from the same product
Cons
- The extension is on the cold start path and consumes function memory, which is a meaningful fraction at small memory settings
- Separate meters for invocations, ingested spans, indexed spans, custom metrics and logs; a high-invocation function fans out across several of them at once
- Edge runtimes are not the strength — the Lambda-centric extension model does not transfer to V8 isolates
Best for: Teams heavily invested in Lambda who want traces, logs and metrics leaving through one extension rather than three separately configured pipelines.
Pricing: Per-function serverless billing combined with the usual separate meters for ingested and indexed spans, custom metrics and log ingest. Invocation volume and telemetry volume are billed on different dimensions, so a high-frequency short function and a low-frequency verbose one cost very differently.
Lumigo

Lumigo is serverless-first rather than serverless-added, which shows in the details. It reconstructs the invocation payload and the AWS service calls around it, which is what you need when a Step Function stalls and the question is what the state machine actually passed to the next state. Deployment is by layer or by automated tracing across an account, and it treats the managed AWS services in the call path as first-class rather than as opaque endpoints.
Pros
- Captures invocation payloads and AWS SDK call detail, so debugging a stalled Step Function does not require adding logging and redeploying
- Understands SQS, SNS and EventBridge context propagation as a primary case rather than an HTTP special case
- Automated instrumentation across an account means new functions are covered without a per-function layer step
- Cold start and init reporting are built around Lambda’s actual phase model rather than mapped onto a generic host-based one
Cons
- Payload capture is exactly the feature that raises data governance questions, and redaction rules become something you must maintain
- Narrow beyond serverless — if your estate is a mix of functions and containers you are buying a second tool for the rest
- Edge runtimes outside the AWS Lambda model are not the focus
Best for: Teams whose architecture is genuinely serverless-native — Lambda, Step Functions, SQS and EventBridge — where the AWS service call path matters more than generic application tracing.
Pricing: Usage-based on traced invocations, with tiers by retention and feature level. Because the meter is invocations rather than hosts or nodes, a high-frequency low-duration function is the expensive shape here.
Honeycomb

Honeycomb suits edge and serverless workloads because its model is wide, high-cardinality events rather than pre-aggregated metrics. That fits a workload where you want to slice by colo, route, function version, memory setting and customer without deciding in advance which of those dimensions you would need. It accepts OTLP over HTTP, which is the only protocol a V8 isolate can speak.
Pros
- OTLP over HTTP ingest works directly from a Cloudflare Worker’s
waitUntil, with no agent and no layer required - High-cardinality querying means colo, function version and cold-start flag can all be dimensions without a series count explosion
- Suits the serverless debugging pattern of asking an unplanned question about a specific slow invocation after the fact
- No per-host or per-function licensing dimension, which matches an environment with no hosts
Cons
- Event-volume billing means a chatty high-invocation function is directly expensive, and sampling becomes something you must design rather than an option
- Fewer AWS-specific serverless features than the serverless-native vendors — no payload reconstruction, no Step Functions execution view
- You build the instrumentation; there is no auto-instrumenting extension doing the work for you
Best for: Edge and serverless teams who debug by asking unplanned high-cardinality questions and are willing to design a sampling strategy to control event volume.
Pricing: Billed on events ingested with tiers by retention, so telemetry volume is the only dimension — invocations, functions and regions are not counted. Sampling is the primary cost control.
Grafana

Grafana accepts OTLP over HTTP directly, which is the only protocol available to you at the edge, and its component split works for the log-based half of the serverless story: Loki takes the structured JSON your function writes to stdout, Tempo takes the spans an extension exports, and both are queryable side by side. For teams already running it for containers, adding functions does not mean a second vendor.
Pros
- OTLP over HTTP ingest works from edge isolates as well as from a Lambda extension
- Loki handles the log-based export path natively, which is the fallback for functions where no layer can be installed
- One backend can serve functions, containers and edge, so a hybrid estate stays on one query surface
- Self-hostable, which matters when the telemetry from an edge function includes request data you cannot send to a vendor
Cons
- No serverless-specific analysis — cold start attribution and init duration splitting are dashboards you build, not features you get
- Requires an extension or a synchronous flush on the Lambda side; there is no vendor layer doing that setup for you
- Several components to operate, and the correlation between Loki logs and Tempo traces depends on label discipline you enforce
Best for: Teams already running Grafana for containers who want serverless and edge telemetry in the same backend rather than adopting a second product.
Pricing: Open source components free to self-host; the managed cloud bills on metric series, log volume and trace volume. Functions add no host or seat dimension, so cost tracks how verbose your handlers are.
SigNoz

SigNoz accepts OTLP over HTTP directly, which is what makes it usable from an edge isolate as well as from a Lambda extension. Because there is no translation layer, the resource attributes the OpenTelemetry Lambda layer sets — function name, version, memory size, cold start flag — arrive as queryable attributes rather than being mapped into a vendor’s host-shaped schema that has no room for them.
Pros
- OTLP over HTTP means both the OTel Lambda layer and an edge Worker can export to it with the same configuration
- Lambda resource attributes survive as first-class queryable fields instead of being flattened into a host model
- Ingest-based billing with no host or function dimension, which matches an architecture with neither
- Self-hostable, so telemetry from functions handling regulated data need not leave your infrastructure
Cons
- No serverless-specific features — cold start attribution, provisioned concurrency utilization and Step Functions views are not there
- You configure the export path yourself; there is no vendor-maintained layer bundling instrumentation and extension
- Self-hosting adds a persistent system to run for a workload you adopted specifically to avoid running persistent systems
Best for: Teams using the OpenTelemetry Lambda layer who want an OTLP-native backend that also accepts exports directly from edge runtimes.
Pricing: Open source and free to self-host; the managed offering bills on data ingested and retention period, with no per-function or per-invocation component.
New Relic

New Relic offers serverless instrumentation and a Lambda layer, worth evaluating if you already use it elsewhere in the estate. The layer follows the same pattern as the others — extension registered during init, telemetry forwarded out of band — and there is also a log-based ingest path via a CloudWatch Logs subscription, which is the option for functions where you cannot add a layer at all. It accepts OTLP, so functions instrumented with the OpenTelemetry layer can report to it without adopting the vendor SDK.
Pros
- Both a layer-based and a log-subscription ingest path, so functions that cannot take a layer are still covered
- Ingest-and-seat billing has no per-function or per-invocation dimension, so a fleet of hundreds of small functions does not multiply the bill
- Accepts OTLP, so the OpenTelemetry Lambda layer is a supported front end and the instrumentation stays portable
- One backend for functions and long-lived services, useful in an estate that is only partly serverless
Cons
- Ingest-based billing sits badly with high-invocation functions, where span volume scales directly with traffic and there is no host count to amortise it against
- The layer adds init work on the cold start path like every other extension, and the log-subscription alternative trades that for log ingest cost
- Serverless-specific depth trails the serverless-native vendors, particularly on payload reconstruction and Step Functions execution views
Best for: Organisations already standardised on New Relic for their containers and VMs who want functions in the same backend without adding a per-invocation meter.
Pricing: Billed on data ingested plus a per-user model with tiers of platform access, with no per-function or per-invocation dimension. Function count is free; how much each invocation emits is what you pay for.
How to choose
Write down your invocation profile. Invocations per month, average duration, cold start rate, and how many functions are queue-triggered versus HTTP-triggered. Serverless pricing is usually per invocation or per ingested byte, so these numbers price the decision.
Measure the layer’s cold start cost on your smallest function, not your biggest.
Test an SQS batch. If the tool shows one parent and four orphans, you know what you are buying.
Check the edge story separately. Do not assume Lambda support implies Workers support.
Decide your flush strategy explicitly: extension wherever a layer works, synchronous flush only where you have measured that the milliseconds do not matter, log-based for edge.
| Lambda extension layer | Edge runtime export | Billing dimension | |
|---|---|---|---|
| Datadog | Vendor layer, traces plus logs plus metrics | Not the focus | Per function plus separate telemetry meters |
| Lumigo | Vendor layer or account-wide auto-tracing | Not the focus | Traced invocations |
| OpenTelemetry SDK (standard, not a product) | Community layer, vendor-neutral | OTLP over HTTP from waitUntil | None |
| Honeycomb | Via OTel layer | OTLP over HTTP | Events ingested |
| Grafana | Via OTel layer, or log-based through Loki | OTLP over HTTP | Data volume |
| SigNoz | Via OTel layer | OTLP over HTTP | Data ingested |
| New Relic | Vendor layer or log subscription | OTLP over HTTP | Data ingested plus seats |
All-in on Lambda and want the least work? A serverless-native vendor gets you further faster than a general-purpose platform with a serverless module bolted on. Running a mix of serverless and containers? Use the OpenTelemetry Lambda layer into whatever backend already serves your services — portability beats the last ten percent of polish and keeps this consistent with the rest of your observability stack.
Frequently asked questions
Why is my Lambda telemetry incomplete or missing entirely?
The flush. The environment freezes when your handler returns, so any background export thread stops mid-flight and its buffer is lost when the environment is reclaimed. Fix it with a Lambda extension that receives data over loopback and flushes on shutdown, or by calling forceFlush() before returning.
Can I run an APM agent on Cloudflare Workers or Vercel Edge?
No binary agent. Edge runtimes are V8 isolates with no filesystem and no native module support, so instrumentation must be JavaScript inside the isolate and export must go over fetch to an OTLP-over-HTTP endpoint. Use the platform’s post-response hook, such as waitUntil, so export does not delay the response.
Why do my SQS-triggered function traces show no parent?
Trace context does not travel in an SQS message unless you put it there. The producer writes traceparent into message attributes and the consumer extracts it. For batch sizes above one, the correct model is span links to each message’s context, and not every tool implements that.
Related reading
- Best APM tools for developers — the full category and how pricing models compare.
- Best APM for Kubernetes workloads — the opposite architecture, where you run an agent but pay per pod.
- Best OpenTelemetry-native observability platforms — backends taking OTLP over HTTP, all edge can send.
- Best APM for Node.js applications — the runtime most serverless functions use.
- Best APM tools for startups — free tiers under invocation-priced workloads.