nemoclaw-benchmark-header.png
  1. Home
  2. Insights

Build Fast, Reliable Agentic AI From Edge To Cloud

10 Sept 2026
Marian Petruk, Milosz Dudek, Alesandro Slepčević

ShareShare the article

In brief

SoftServe's R&D team built and benchmarked a complete hybrid edge-cloud AI pipeline on Akamai's GPU cloud services, using NVIDIA NemoClaw™ open blueprints to power the agent layer. The key findings appear below. The full architecture breakdown and benchmark data live in a report at the end of this post.

  • ~22 ms average server-side inference latency on Akamai's NVIDIA RTX 4000 Ada, and ~28.5 ms end-to-end round trip in the full hybrid pipeline
  • The NVIDIA RTX PRO 6000 Blackwell Server Edition delivers the highest measured throughput (638K images/hour)
  • NVIDIA OpenShell™ —the secure Runtime included in NemoClaw blueprints —enforced policy-governed agent behavior throughout, providing the architectural foundation for sensitive data isolation in production deployments

AI agents need to act, but data must be protected

Enterprises are under growing pressure to deploy autonomous AI agents that can reason across their data, take actions in their systems, and operate continuously without human bottlenecks. The use cases, such as clinical documentation, fraud triage, predictive maintenance, supply chain intelligence, and smarter customer operations, are clear.

What's less clear is how to deploy agents safely. Most enterprise data, patient records, financial transactions, proprietary operational telemetry, customer PII, cannot be routed through a public AI API. Most agent frameworks don't provide the runtime policy controls, audit trails, or sandboxed execution environments that regulated industries require.

The result is a pattern we see repeatedly. Promising AI pilots that work in a controlled environment and can't move to production because the governance story doesn't hold up.

Compliance teams, CISOs, and customers don't spend much time asking whether to deploy AI agents anymore. They ask whether they can trust how the agent behaves once it's running.
Alesandro Slepčević
Akamai Principal Cloud Architect

Architecture built for governance

Our R&D team set out to design and benchmark a production-pattern AI inference pipeline that addresses the governance problem directly. The architecture has three layers.

  1. Edge device.

    A laptop, local PC, or smart glasses connected to the cloud.

  2. Akamai Cloud.

    The edge device's output goes to an Akamai NVIDIA GPU instance running a DINO (self-DIstillation with NO labels) Matching visual classification model. This is where the compute-intensive work happens — classification, similarity matching, feature extraction — on Akamai's NVIDIA AI infrastructure.

  3. NVIDIA NemoClaw™ agent layer.

    The classification results feed into a NemoClaw™ -powered agent that applies business logic and responds by text or voice. It runs under NVIDIA OpenShell™ runtime controls: sandboxed access, policy-governed model routing, and an immutable log of every model call, data access, and action taken.

The pipeline is genuinely hybrid. Latency-sensitive detection happens at the edge. Compute-intensive classification happens in the cloud. The agent layer, the part that reasons and acts, is governed throughout by OpenShell policy controls. Sensitive data need not be exposed to a public API at any stage.

nemoclaw-benchmark-image.png

Our benchmarks

We ran two benchmark tracks on Akamai's GPU infrastructure in Frankfurt: an application-facing latency test on the NVIDIA RTX 4000 Ada, and a raw throughput test across the RTX 4000 Ada and RTX PRO 6000 Blackwell. All figures below come from SoftServe/Akamai R&D testing, May–June 2026.

 

SoftServe/Akamai POC, 2026

~22 ms

Average server-side inference latency, DINO Matching on Akamai RTX 4000 Ada

~28.5 ms

Akamai-to-Akamai round-trip latency (2.6 ms client-to-cloud, 22.7 ms server, 3.2 ms cloud-to-client)

110K

Peak raw GPU throughput, RTX 4000 Ada at optimal RT-DETR batch config

638K

Peak raw GPU throughput, RTX PRO 6000 Blackwell at optimal RT-DETR batch config

Two patterns stood out. Latency is mostly a function of distance: on a public-internet route, network transfer dominates the round trip, but once the client and GPU sit close together on Akamai's network, processing time becomes the bottleneck instead. And the faster GPU isn't automatically the cheaper one — Blackwell's throughput advantage tips the cost math in North America, while Ada wins on price per image in Europe.

Auditable, fast compliance

The specific POC scenario we used, a visual AI pipeline with a voice-interactive NemoClaw™ agent, is one application of a more general architecture pattern. The same design applies wherever an organization needs AI agents that act on sensitive data at speed, with policy controls and audit trails, without routing that data through a public cloud AI service.

The NemoClaw open blueprint stack — and specifically NVIDIA OpenShell — is what makes that governance operational and fast. It does not remove the organization's own compliance obligations. Meeting requirements like HIPAA, GDPR, and other PHI/PII rules end-to-end still requires building those controls into the security blueprint, with a self-hosted or local LLM as the strongest option for keeping regulated data fully in-house.

This is already relevant across healthcare, financial services, manufacturing, and technology — anywhere agents touch regulated data or systems of record. The whitepaper walks through what that looks like in each industry.

The role of NVIDIA NemoClaw

NVIDIA NemoClaw is a collection of open blueprints for building specialized enterprise agents on a stack organizations can own, customize, and run anywhere.. NVIDIA OpenShell— the secure runtime included in NemoClaw blueprints—defines what models can be called, what data they can access, what actions agents can take, and what gets logged, before any inference call is made.

NemoClaw™ 's open-source reference stack means SoftServe can implement it as a delivery pattern, the same governance architecture applied to a clinical documentation agent, a supply chain intelligence agent, or a fraud investigation workflow, rather than re-engineering the controls from scratch for each use case.

Get the Full Benchmark

This article covers the headline numbers. The benchmark report covers the complete architecture design, the full pipeline component breakdown, the methodology behind every figure above, detailed performance and regional cost-efficiency tables across GPU instances and batch sizes, and a guide to applying this pattern across regulated industry workloads.

Governing AI Agents in Regulated Industries

ShareShare the article

Authors

Marian Petruk

Marian Petruk

SoftServe R&D Cluster L ead (Effi cient AI)

LinkedInMore from this author
Milosz Dudek

Milosz Dudek

SoftServe Middle R&D Engineer

LinkedInMore from this author
Alesandro Slepčević

Alesandro Slepčević

Akamai Principal Cloud Architect

LinkedInMore from this author

Don't want to miss a thing?

Subscribe to get expert insights, in-depth research, and the latest updates from our team.
SUBSCRIBE

Our Insights

white paper
Generic agentic AI

Generative and Agentic AI Trends for 2026

Explore more
white paper

The Economics of AI: Cost Dynamics and New Value Streams

Explore more
white paper
AI Embodied

Humanoids: AI Embodied in Physical Form

Explore more
white paper

Cloud Modernization: The AI Catalyst No One Can Ignore

Explore more
guide

AI and the Cost of Certainty

Explore more
Softserve company logo
KarriereSearchContact Us

Hot Links

  • Home
  • Branchen
  • Services und Kompetenzen
  • Ressourcen
  • Presse
  • Über Uns
  • Kontakt
  • Karriere
  • Subscribe to Updates

Kontakte

  • Hauptsitz Austin

    201 W 5th Street Suite 1550 Austin, TX 78701

    +1-512-516-8880

    Toll Free:
    +1-866-687-3588

  • Privacy Notice - 
  • Terms and Conditions - 
  • Information Security - 
  • Sitemap - 
  • Search - 
  • Accessibility statement - 
  • LInkdn Link with Icon
  • Youtube Link with icon
  • Facebook Link with Icon
  • Instagram Link with icon

© Copyright 2026 SoftServe Inc.

hero-icon.svg
TikTok Link with icon
  • Twitter Link with icon
  • Soundcloud Link wiht icon
  • Bluesky Link with icon