{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Rangoon Insights",
  "home_page_url": "https://rangoon.ai/insights/",
  "feed_url": "https://rangoon.ai/feed.json",
  "description": "AI news, analysis, and practical technical guides.",
  "language": "en",
  "authors": [
    {
      "name": "Rangoon Editorial",
      "url": "https://rangoon.ai/insights/editorial/"
    }
  ],
  "items": [
    {
      "id": "https://rangoon.ai/insights/anthropic-claude-sonnet-5-5-release-engineering/",
      "url": "https://rangoon.ai/insights/anthropic-claude-sonnet-5-5-release-engineering/",
      "title": "Claude Sonnet 5.5 makes model selection a release-engineering task",
      "summary": "Anthropic released Claude Sonnet 5.5 on September 28. Its arrival is a reminder to treat model upgrades as measurable, reversible application changes.",
      "content_text": "A new model is an input to a running product\n\nAnthropic announced Claude Sonnet 5.5 on September 28 as the second release in its Claude 5.5 family. The company presents it as a faster, lower-cost complement to Claude Opus 5.5 for well-scoped work, bug fixing, and document-oriented tasks. That positioning is useful for planning, but it is not a substitute for testing a product's own tasks.\n\nA model upgrade changes more than text quality. It can change how often an agent chooses a tool, how it handles ambiguous input, the shape of structured output, the amount of reasoning it produces, and the timing of a multi-step workflow. Each of those changes can expose assumptions in prompts, schemas, UI states, budgets, or downstream parsers.\n\nThe useful framing is release engineering. A provider announcement identifies a new candidate. The application team decides whether that candidate is compatible with its product, where it is allowed to run, and what evidence is sufficient to move it from evaluation to a user-facing path.\n\nTest representative work instead of chasing a headline metric\n\nAnthropic publishes its own capability, speed, and cost comparisons for Sonnet 5.5. Those figures can help a team decide what to investigate, but they are provider results under provider-selected conditions. They cannot establish the behavior of a specific codebase, support workflow, or regulated process.\n\nA practical evaluation set contains the work users actually ask for: ordinary requests, incomplete requests, long documents, formatting-sensitive outputs, tool calls that should be declined, and known failure cases. Keep the inputs and expected review criteria stable across models. Then compare success rate, repair work, latency distribution, token use, and the cases where a human reviewer would reverse the result.\n\nThe same discipline applies to model controls. If an application depends on a certain effort setting, structured-output mode, or tool behavior, treat it as a tested configuration rather than an incidental default. A prompt that appears portable can still produce different operational behavior after a provider upgrade.\n\nPin the candidate model name and configuration in an evaluation environment.\n\nRun a versioned set of representative tasks, including failure and recovery cases.\n\nDefine a rollback signal before exposing the upgrade to a broader audience.\n\nModel lifecycle information belongs in the deployment record\n\nAnthropic's platform documentation maintains a separate view of model status and retirement information. That is an important operational source because an application can remain technically correct while relying on a model that is approaching a provider lifecycle change.\n\nA mature model record names the provider identifier, application configuration, prompt or tool contract version, evaluation set, traffic policy, and owner. It also records the fallback behavior if the provider changes availability, returns an error, or removes an older model from service. This is ordinary dependency management, applied to a probabilistic component.\n\nSonnet 5.5 is therefore most useful as a concrete reason to improve that record. The durable outcome is not a claim that one model wins everywhere. It is a repeatable process for selecting a model based on observed fit, cost boundaries, and recoverable operations.\n\nSources\nAnthropic: Introducing Claude Sonnet 5.5: https://www.anthropic.com/claude-sonnet-5-5\nAnthropic Newsroom: https://www.anthropic.com/news\nClaude Platform Docs: Model deprecations: https://docs.anthropic.com/en/docs/about-claude/model-deprecations",
      "image": "https://rangoon.ai/assets/editorial/article-models-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "models",
        "News analysis"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/docker-cloud-sandboxes-agent-isolation/",
      "url": "https://rangoon.ai/insights/docker-cloud-sandboxes-agent-isolation/",
      "title": "Cloud Sandboxes move agent isolation beyond the laptop",
      "summary": "Docker's Cloud Sandboxes extend its microVM model to managed compute. Teams now need to test portability, spend controls, and workflow handoffs.",
      "content_text": "The cloud is becoming another sandbox location\n\nDocker announced Cloud Sandboxes on September 24, then used an October 1 event recap to describe the same direction in operational terms: a developer can begin with the sandbox workflow locally and continue it on Docker-managed compute. Docker describes the cloud environment as a microVM-based sandbox with its own kernel and Docker daemon, rather than as a shared container dropped into an ordinary host.\n\nThat is a meaningful shift for agent work. A local sandbox is useful for quick iteration, but long-running tasks, controlled egress, and team handoffs often need capacity that does not live on one laptop. A common sandbox model can reduce the number of environment-specific workarounds a tool author has to maintain.\n\nThe event itself is Docker's product announcement, not a claim about a universal deployment standard. Its practical value will depend on how teams apply it to their own images, workloads, retention rules, and review process.\n\nPortable files still meet local services\n\nDocker says its move workflow captures filesystem state and recreates that state in the destination. That makes a useful boundary visible: project files and tool state can travel, while environment credentials, templates, and network policies are managed separately for local and cloud use.\n\nThat separation is useful for everyday development. A task snapshot may contain a branch, dependencies, generated files, and test artifacts. The destination still has to supply its own credentials, egress settings, image cache, capacity limits, and storage rules. Treating these as visible environment services prevents a local convenience from becoming an unexplained cloud dependency.\n\nThe practical handoff question is whether a developer can reproduce the work without dragging along accidental state. A good exercise records the image, working tree, inputs, service dependencies, and expected output, then compares the local result with the cloud result. Differences in network access, cache warmth, or available compute should be observable rather than surprising.\n\nTreat code and working files as portable task state.\n\nTreat credentials, egress rules, and execution permissions as destination-specific controls.\n\nMeasure cold starts, idle time, and retained storage before treating a remote sandbox as a default workstation.\n\nThe useful question is what survives a handoff\n\nDocker's Sandbox Kit specification frames a sandbox as an OCI image plus a small set of runtime conventions. That makes a composable base for experiments, but it does not answer every operational question around an autonomous workload. Operators still need to decide image ownership, dependency update cadence, data retention, and what happens when a task needs more time or capacity than expected.\n\nTeams evaluating cloud sandboxes can begin with a narrow, measurable exercise: run the same contained task locally and remotely, then compare setup time, output, logs, network behavior, and total resource use. The result is more useful than a generic portability claim because it exposes which parts of the workflow are actually portable.\n\nDocker's announcement is a signal that isolation and portability are becoming part of the ordinary developer path. The remaining work is operational: choose the right workloads, set practical budgets, and make the transition between laptop and managed compute easy to understand.\n\nSources\nDocker: Cloud Sandboxes recap from WeAreDevelopers: https://www.docker.com/blog/docker-cloud-sandboxes-wearedevelopers-recap/\nDocker: Introducing Cloud Sandboxes: https://www.docker.com/blog/introducing-cloud-sandboxes-start-on-your-laptop-finish-in-the-cloud/\nDocker: Sandbox Kit Specification: https://www.docker.com/blog/docker-sandbox-kit-spec/",
      "image": "https://rangoon.ai/assets/editorial/article-containers-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "infrastructure",
        "News analysis"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/google-gemini-3-8-flash-ga-migration/",
      "url": "https://rangoon.ai/insights/google-gemini-3-8-flash-ga-migration/",
      "title": "Gemini 3.8 Flash turns model migration into an API operations exercise",
      "summary": "Google released Gemini 3.8 Flash on September 2. The GA milestone puts version choice, lifecycle tracking, and workload testing back in focus.",
      "content_text": "General availability is a useful starting signal\n\nGoogle's Gemini API release notes list Gemini 3.8 Flash as generally available on September 2. Google describes the model as aimed at long-horizon software engineering, autonomous agents, and complex enterprise workflows. That tells developers where Google expects it to be useful, while leaving each application responsible for deciding whether the behavior fits its own product.\n\nGeneral availability changes the planning conversation because it gives teams a named model and published lifecycle entry to evaluate. It does not make an upgrade automatic. The model may be called through several APIs, combined with tools, constrained by structured-output requirements, or used in a latency-sensitive interaction where a small behavior difference matters more than a broad capability label.\n\nThe first engineering task is to identify every place the previous model is selected. That includes server configuration, worker jobs, background queues, evaluation scripts, support tools, and any provider abstraction that may have a hidden default.\n\nMigration should be treated as an application change\n\nGoogle's latest-model guidance describes Gemini 3.8 Flash in terms of software work, agents, and enterprise workflows. Those broad categories are not a test plan. A coding assistant may depend on file-edit behavior; a support assistant may depend on stable classifications; a workflow agent may depend on reliable tool arguments and predictable retries.\n\nStart by replaying a fixed sample of production-like inputs in an isolated environment. Compare output schema validity, tool selection, error recovery, latency, token use, and the amount of human cleanup required. When the model is part of a chain, test the entire chain rather than judging the first response in isolation.\n\nThen use a bounded rollout with direct observability. Track which model handled each request, keep a known fallback, and set a threshold that pauses the rollout when quality or operations regress. That approach turns an API upgrade into a reversible product change.\n\nAudit explicit and implicit model selection before replacing a default.\n\nTest structured output, tool calls, retry paths, and long-running tasks with the real application contract.\n\nRoll out gradually only after fallback and request-level model telemetry are in place.\n\nLifecycle information is operational data\n\nGoogle's Gemini API deprecations page lists a release date for Gemini 3.8 Flash and no announced shutdown date at the time of this article. That is more useful than treating model names as permanent infrastructure: it makes clear that a provider maintains a lifecycle and that an application needs to watch it.\n\nA dependency inventory for model-backed features should include the exact endpoint, API version, region or service configuration where applicable, prompt and tool contract, owner, and last evaluation date. It should also say how the product responds when the provider returns a capacity error or a model becomes unavailable.\n\nThis approach avoids both extremes. Teams do not need to postpone every upgrade until a perfect comparison exists, and they do not need to accept a new default without evidence. They can make an explicit, measured choice and retain enough information to revisit it when the provider's lifecycle changes.\n\nSources\nGoogle AI for Developers: Gemini API release notes: https://ai.google.dev/gemini-api/docs/changelog\nGoogle AI for Developers: What's new in Gemini 3.8 Flash: https://ai.google.dev/gemini-api/docs/latest-model\nGoogle AI for Developers: Gemini deprecations: https://ai.google.dev/gemini-api/docs/deprecations",
      "image": "https://rangoon.ai/assets/editorial/article-models-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "models",
        "News analysis"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/mcp-stateless-core-tool-infrastructure/",
      "url": "https://rangoon.ai/insights/mcp-stateless-core-tool-infrastructure/",
      "title": "MCP's stateless core changes the shape of tool infrastructure",
      "summary": "MCP's July specification moved its core to stateless operation. The change makes ordinary web infrastructure a better fit for tool-service migration.",
      "content_text": "MCP moves closer to ordinary web operations\n\nOn July 28, the Model Context Protocol project published a new specification revision with a stateless core. Its release notes describe the shift away from protocol-level sessions and toward request metadata, header routing, and cacheable lists. The supporting SDK announcement had already explained the operational consequence: servers can fit more naturally behind ordinary load balancers without relying on sticky session state.\n\nThat is an infrastructure change with practical service consequences. Tool providers can use familiar routing, observability, and capacity patterns instead of treating every MCP connection as a special long-lived control channel. A tool service can sit behind standard load balancing, provided its application state is stored and recovered deliberately.\n\nThe revision is a published protocol milestone, not proof that every existing server has migrated. Teams should identify the specification version their clients and servers negotiate, then test actual behavior under the routing, retry, and failure modes they operate.\n\nStateless transport does not erase application state\n\nRemoving a protocol session does not remove context. It moves responsibility for context into the application and requests themselves: caller identity, tenant or workspace, selected tool, timeouts, correlation identifiers, and any task state that must survive a retry. A service that relied on a sticky connection must make those dependencies explicit.\n\nThat can improve operational clarity. A gateway can route a request by ordinary HTTP metadata, while the service stores durable work in its own database or queue. The migration risk is assuming that the absence of a protocol session automatically makes an interaction safe to retry, cache, or distribute across replicas.\n\nThe extension framework creates a related compatibility task. An extension can add features to an exchange, but clients and servers need a clear fallback when a peer does not recognize it. Capability negotiation should be logged and covered by contract tests.\n\nNegotiate protocol and extension support explicitly.\n\nCarry correlation and tenant context in a documented request path.\n\nTest retries, caching, cancellation, and load balancing with a real deployment topology.\n\nA migration plan should start at the service edge\n\nThe release documentation describes new protocol mechanisms such as Multi Round-Trip Requests and a formal extension framework. Those features can make richer interactions possible between a client and tool server, but they also create more code paths to observe, time out, and test under partial failure.\n\nA sound migration starts at the service edge. List every client version, map session assumptions, decide where durable task state lives, and introduce compatibility tests before changing the production routing model. Then measure connection reuse, cache behavior, error handling, and rollback paths with representative traffic.\n\nDocumenting these decisions makes incident response faster because operators can distinguish protocol compatibility failures from ordinary service outages.\n\nThe broader lesson is architectural rather than promotional: a protocol can make service boundaries cleaner, but only an application migration plan makes those boundaries dependable in production.\n\nSources\nModel Context Protocol: The 2026-07-28 Specification: https://blog.modelcontextprotocol.io/posts/2026-07-28/\nModel Context Protocol: Beta SDKs for the 2026-07-28 specification: https://blog.modelcontextprotocol.io/posts/sdk-betas-2026-07-28/\nModel Context Protocol 2026-07-28 changelog: https://github.com/modelcontextprotocol/modelcontextprotocol/blob/main/docs/specification/2026-07-28/changelog.mdx",
      "image": "https://rangoon.ai/assets/editorial/article-models-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "infrastructure",
        "News analysis"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/nvidia-openshell-nemoclaw-boundaries/",
      "url": "https://rangoon.ai/insights/nvidia-openshell-nemoclaw-boundaries/",
      "title": "OpenShell and NemoClaw draw a cleaner line around agents",
      "summary": "NVIDIA's OpenShell and NemoClaw split the agent harness from its runtime controls. That separation gives security reviews clearer seams to inspect.",
      "content_text": "The runtime and the agent stack are separate products\n\nOn March 16, NVIDIA presented OpenShell as an open-source runtime for running autonomous agents inside isolated sandboxes, with a policy engine and privacy routing around the workload. The same announcement described NemoClaw as a stack that pairs an agent harness and models with that runtime for always-on assistant scenarios.\n\nNVIDIA's current documentation makes the boundary more explicit. OpenShell supplies the general-purpose sandbox and policy platform. NemoClaw adds the managed workflow around a particular agent experience: onboarding, lifecycle handling, configuration blueprints, and related operations. A lower-level OpenShell interface remains available for teams that need direct runtime control.\n\nThat division is valuable because it avoids treating an agent's prompt, tool loop, and operating boundary as one indistinguishable feature. Each layer can change at a different pace and can be reviewed against different operational concerns.\n\nThe stack still has an operational scope\n\nNVIDIA describes OpenShell as governing execution, visible resources, and inference routing. Those controls can keep credentials outside an agent process, constrain destinations, and make the runtime boundary easier to inspect. They do not erase the need to understand the agent harness, the model endpoint, the host operating system, or the application services outside the sandbox.\n\nThe practical review question is simpler than a broad safety claim: which component creates the sandbox, which component stores configuration, which process receives credentials, and which layer observes failures? Teams should be able to answer that from deployed configuration and logs, rather than from a product diagram alone.\n\nFor sensitive external effects, a sandbox and a network rule are useful controls but not a substitute for the organization's own identity and change-management process. That is one checkpoint in a larger operating model, not the purpose of every agent workload.\n\nHarness: plans work, maintains task context, and requests tools.\n\nRuntime: isolates processes, routes access, and applies execution constraints.\n\nOperations: owns image updates, service accounts, logs, recovery, and the surrounding change process.\n\nThe architectural lesson travels beyond one vendor stack\n\nNemoClaw's documented use of a versioned blueprint is a useful reminder that repeatable agent environments need explicit configuration. Images, policies, and inference profiles should be inspectable inputs to a deployment, not hidden defaults inferred after a task has started.\n\nThat also makes change review more concrete: an operator can compare a declared profile with the environment that actually ran, instead of relying on a general statement that the agent was sandboxed.\n\nFor teams building with multiple runtimes, the useful portable contract is smaller than a full platform replacement. Define how a harness receives configuration, how it selects an inference endpoint, how it reports health, and how a stopped or failed sandbox is recovered. Then a runtime such as OpenShell can be evaluated as one execution option instead of a catch-all answer to agent operations.\n\nThe key lesson from NVIDIA's architecture is specificity. A published sandbox boundary is useful when its image, process, network, and configuration boundaries are clear enough for engineers to test under their own workloads.\n\nSources\nNVIDIA Technical Blog: Run Autonomous, Self-Evolving Agents More Safely with NVIDIA OpenShell: https://developer.nvidia.com/blog/run-autonomous-self-evolving-agents-more-safely-with-nvidia-openshell/\nNVIDIA NemoClaw documentation: Overview: https://docs.nvidia.com/nemoclaw/latest/about/overview.html\nNVIDIA NemoClaw documentation: Choose Between NemoClaw and OpenShell CLIs: https://docs.nvidia.com/nemoclaw/latest/user-guide/openclaw/reference/cli-selection-guide",
      "image": "https://rangoon.ai/assets/editorial/article-agents-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "agents",
        "News analysis"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/a2a-v1-stable-agent-handoffs/",
      "url": "https://rangoon.ai/insights/a2a-v1-stable-agent-handoffs/",
      "title": "A2A v1.0 makes agent handoffs easier to inspect",
      "summary": "A2A v1.0 stabilized an open protocol for agent-to-agent work. The milestone makes interfaces and version migration more concrete for builders.",
      "content_text": "A stable protocol gives delegation a clearer boundary\n\nThe A2A Protocol community announced version 1.0 on March 12 as the first stable, production-ready release of its open standard for communication between AI agents. The announcement emphasizes a common semantic model, multiple protocol bindings, and version negotiation so systems can describe compatible ways to work together across framework and vendor boundaries.\n\nFor an engineering team, that matters because agent handoffs have often been embedded inside one framework's internal conventions. A shared protocol makes it more practical to expose an agent interface, advertise what it can handle, and exchange work without first rebuilding every peer around the same runtime.\n\nThe milestone is about interoperability, not a blanket claim that connected agents are safe or equivalent. A2A gives systems a language for a handoff; it does not decide whether a handoff should be accepted in a given organization.\n\nA version migration is more than a schema update\n\nA2A documentation shows that an Agent Card can describe supported interfaces and security requirements. That gives a client a structured discovery record, but it also creates migration work: clients must choose an advertised interface, send the negotiated protocol version, and handle a peer that does not support a newer capability.\n\nTask behavior deserves the same care. An agent can return messages, artifacts, status transitions, and a request for later work. Implementers need to decide which of those become durable application records, how retries are deduplicated, and how cancellation behaves when a receiving service is slow or temporarily unavailable.\n\nThe protocol's value is that these problems appear at a clear boundary. Teams can capture the selected Agent Card, interface, version, request identifier, and resulting task state in normal service telemetry, then reproduce failures without inferring them from model conversation history.\n\nDiscover an agent interface before selecting it.\n\nTest supported versions and bindings against a real peer before a broad rollout.\n\nDefine retries, cancellation, artifact storage, and terminal task states in the receiving service.\n\nStart with observable, reversible handoffs\n\nThe A2A project continued with a 1.0.1 maintenance release in May, a reminder that a stable protocol still evolves through implementation feedback. Adopting it should include version handling, compatibility tests, and a way to see which interface and binding were selected for every request.\n\nA practical first use is a constrained handoff between two distinct services: one prepares a structured research request, another returns an artifact, and the application records the full task lifecycle. The test is successful when a developer can replay a failed or interrupted exchange from the task data and service logs.\n\nSmall integration exercises also reveal assumptions about artifact size, time limits, and retry ownership before those assumptions become production incidents.\n\nFor teams with existing agent tooling, the first payoff is often modest but real: an explicit interface contract makes it easier to replace a peer, add a second implementation, or expose task state to ordinary observability systems.\n\nSources\nA2A Protocol: A2A Protocol Ships v1.0: https://a2a-protocol.org/dev/blog/2026/03/12/a2a-protocol-ships-v10-production-ready-standard-for-agent-to-agent-communication/\nA2A Protocol blog archive: https://a2a-protocol.org/latest/blog/\na2aproject/A2A releases: https://github.com/a2aproject/A2A/releases",
      "image": "https://rangoon.ai/assets/editorial/article-agents-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "agents",
        "News analysis"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/docker-model-runner-local-inference/",
      "url": "https://rangoon.ai/insights/docker-model-runner-local-inference/",
      "title": "Docker Model Runner and the local inference boundary",
      "summary": "A practical guide to local model serving with Docker Model Runner, resource checks, API exposure, isolation, and evaluation discipline in practice.",
      "content_text": "What Docker Model Runner changes\n\nDocker Model Runner gives a Docker-managed place to pull, cache, and serve models from OCI registries or Hugging Face. Its useful abstraction is model distribution plus an inference engine: a team can move a model artifact through familiar registry controls, then expose a local serving interface to an application.\n\nThe important boundary is inference. A model endpoint produces tokens or other model results; it does not decide whether a user or agent may read a repository, send a message, or mutate a business system. Keep inference selection separate from runtime tool policy and business-operation approval.\n\nChoose engine and hardware together\n\nDocker documents llama.cpp, vLLM, and Diffusers as distinct engine choices. llama.cpp is oriented toward efficient local development and GGUF models; vLLM targets throughput-oriented serving with Safetensors; Diffusers covers image generation. Hardware and operating system determine which combination is credible, so a compatibility record should name the engine, model format, accelerator, driver, architecture, context limit, and quantization.\n\nDo not compare a laptop CPU run with a GPU server run as if they were one deployment class. Capture time to first token, generation throughput, memory pressure, context length, concurrency, cold-start behavior, and output/tool-call conformance. A benchmark is evidence for a specific workload and environment, not a general release guarantee.\n\nTreat reachability as a security decision\n\nDocker’s documentation states that the Model Runner API is unauthenticated and that any client able to reach it can send inference requests and manage models. That makes network placement a first-class control. Keep the endpoint on an owner-controlled boundary, inspect container-to-host reachability, and decide which workloads may submit prompts or load model artifacts.\n\nA container or sandbox can reduce blast radius, but it does not establish identity, approval, or business authorization. Store credentials by reference, restrict egress, separate development and sensitive workloads, and record who can change the model catalog. Never pass secrets or customer data to a local endpoint merely because it is on a private network.\n\nTurn observations into a deployment decision\n\nA useful evaluation asks four questions: can the model produce the required structured output, can the runtime preserve the declared tool boundary, can the serving environment meet data and latency requirements, and can the team reproduce the chosen configuration? Record model identity, engine, endpoint exposure, hardware, limits, data route, and rollback path as one decision packet.\n\nIn a governed agent stack, these observations can populate a versioned model profile. The profile informs runtime selection and compatibility review; it does not authorize a business side effect. Exact tool authority still belongs to the application’s policy and approval path.\n\nSources\nDocker Model Runner documentation: https://docs.docker.com/ai/model-runner/",
      "image": "https://rangoon.ai/assets/editorial/article-containers-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "infrastructure",
        "Guide"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/mcp-tools-authorization-boundaries/",
      "url": "https://rangoon.ai/insights/mcp-tools-authorization-boundaries/",
      "title": "MCP tools: useful transport, separate authority",
      "summary": "Understand MCP hosts, clients, servers, tools, consent, and Rangoon’s exact-action authorization boundary before connecting agent systems safely.",
      "content_text": "Start with the three MCP roles\n\nThe MCP specification separates the host application that initiates connections, the client connector inside that host, and the server that offers resources, prompts, or tools. The protocol uses JSON-RPC messages and capability negotiation to make integrations composable. This is a clean transport and discovery model, especially when an agent application needs to expose several external systems through one interaction pattern.\n\nThe roles describe communication responsibilities. They do not say that a server may approve a request, that a client may widen a user’s permissions, or that a tool call is an acceptable business operation. Keep those questions in the application and authority layer around MCP.\n\nReview tools as untrusted capabilities\n\nMCP’s own safety guidance treats tools as potentially arbitrary code execution paths and calls for explicit user consent before invocation. A tool name or description is therefore not enough to establish behavior. Review the server identity, source, version, declared inputs, output shape, data classes, network access, side effects, retry semantics, and expected receipt before an agent can propose using it.\n\nThe practical test is to compare the declared operation with the actual target and payload. A tool that says “update” still needs a resource identifier, scope, expected mutation, idempotency key, and policy context. Treat server-provided annotations and returned data as evidence to validate, not instructions that override the caller or repository policy.\n\nPut exact action authority around the tool call\n\nRangoon models an MCP tool as a connector or adapter operation. The runtime can select a tool and produce a versioned action packet, but the packet still needs authenticated identity, project and environment scope, target resource, exact parameters, risk, policy input, approval where required, expiry, and idempotency. LNSAT can then create server-side authorization, consume a one-time capability, and bind the result to a receipt or outcome-unknown state.\n\nMCP transport, OAuth context, a connector credential, and a successful capability negotiation do not authorize the exact side effect. They are inputs to the decision. This separation keeps a read-only tool, a write tool, and a privileged tool visibly different even when they share one protocol.\n\nA connection review checklist\n\nBefore adding an MCP server, record the owner and source, pin its version, inspect its declared tools, and test malformed, ambiguous, oversized, and adversarial inputs. Confirm which data can leave the host, which credentials are referenced, and which actions require a human decision. Test disconnect, timeout, duplicate request, server replacement, and stale approval paths.\n\nAfter the review, keep installation, enablement, credential assignment, agent permission, policy evaluation, approval, authorization, execution, and audit as separate records. Rangoon’s launch architecture provides the capability and evidence model; it does not claim that this public site exposes an MCP execution endpoint.\n\nSources\nModel Context Protocol specification: https://modelcontextprotocol.io/specification/latest",
      "image": "https://rangoon.ai/assets/editorial/article-agents-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "agents",
        "Guide"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/ollama-vllm-inference-tradeoffs/",
      "url": "https://rangoon.ai/insights/ollama-vllm-inference-tradeoffs/",
      "title": "Ollama and vLLM: choosing an inference home",
      "summary": "Evaluate Ollama and vLLM by workload, hardware, API behavior, observability, and governance instead of treating local inference as one mode.",
      "content_text": "Two useful serving shapes\n\nOllama documents local and cloud endpoints, including an OpenAI-compatible surface, with local requests using a local server and cloud requests using an API key. That shape is convenient for a developer workstation, a small team lab, or a controlled prototype where a compact operational footprint matters.\n\nvLLM presents a broader serving toolkit: offline and online inference, an OpenAI-compatible server, batching and parallel deployment paths, metrics, Docker and Kubernetes deployment references, and model-specific configuration. That shape fits teams that need to reason about throughput, concurrency, GPU placement, or a shared serving tier. The distinction is workload shape, not a universal quality ranking.\n\nCompare the tradeoffs that affect agents\n\nFor Ollama, measure setup friction, model availability, local storage, cold-start time, context behavior, API feature coverage, and how local versus cloud requests are separated. For vLLM, measure GPU memory, model support, batching, parallelism, queueing, observability, rollout complexity, and the operational cost of keeping a serving fleet healthy.\n\nAn OpenAI-compatible endpoint is a transport compatibility signal. It does not guarantee equal tokenizer behavior, tool-call formatting, structured-output reliability, context limits, refusal behavior, latency, or retention. Record those differences in a model profile so a runtime can select a target with evidence instead of guessing from a URL shape.\n\nBuild a repeatable evaluation\n\nUse a fixed, redacted workload set that covers ordinary generation, long context, structured output, tool selection, refusal, interruption, and malformed input. Record model identifier, revision, serving stack, hardware, context size, sampling settings, concurrency, queue time, time to first token, token throughput, error class, and data-routing policy.\n\nRepeat the same workload on the local workstation and the shared server. Compare quality and tool-call conformance separately from latency and cost. Include restart, model eviction, provider outage, partial stream, duplicate request, and rollback tests. A model-serving result becomes useful only when another team can reproduce its environment and understand which constraints produced it.\n\nUse serving evidence in an agent stack\n\nA useful model profile binds the exact model, provider or endpoint class, context and output limits, modalities, tool-call representation, usage evidence, data restrictions, and fallback rules. The runtime consumes that profile for selection and compatibility checks; the connector layer remains responsible for external operations.\n\nServing availability never grants execution authority. A model can propose a tool call, and a runtime can format it correctly, while the surrounding authority layer still evaluates the exact packet, approval, target, scope, expiry, and connector operation. Protect endpoints, keep inference logs and prompts within the selected data policy, and create a new review record when the serving stack changes.\n\nSources\nOllama API introduction: https://docs.ollama.com/api/introduction\nvLLM documentation: https://docs.vllm.ai/en/latest/",
      "image": "https://rangoon.ai/assets/editorial/article-models-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "models",
        "Guide"
      ]
    },
    {
      "id": "https://rangoon.ai/insights/opa-cedar-identity-exact-authority/",
      "url": "https://rangoon.ai/insights/opa-cedar-identity-exact-authority/",
      "title": "OPA, Cedar, identity, and the exact action",
      "summary": "Compare policy engines and identity providers with execution authority, approvals, receipts, and the narrow scope of one consequential action.",
      "content_text": "Separate policy decision from enforcement\n\nOPA describes a general-purpose policy engine that accepts structured input, evaluates policy and data, and returns structured decision output. That makes it useful for expressing reusable rules across services, infrastructure, and workflows. The application that asks OPA still owns the request lifecycle, enforcement point, error handling, and evidence binding.\n\nCedar describes a policy language and authorization engine that evaluates principals, actions, resources, entities, and request context. Its schema validates policy structure, while the application supplies the entities and request. In both cases, a policy answer is one decision input. It does not itself consume a one-time capability, perform a connector operation, or prove the resulting external effect.\n\nIdentity narrows context; it does not widen authority\n\nAn identity provider can authenticate a person, service, or workload and attach groups, claims, scopes, session state, or device facts. Those claims help the policy layer understand who is asking and under what context. They remain subject to freshness, audience, issuer, resource, and session checks.\n\nAn OAuth scope is delegated access language, not a blanket grant to an AI agent. A valid token can still be too broad, too old, aimed at the wrong resource, or outside the user’s approved task. Rangoon records identity as part of a canonical packet, then evaluates the exact operation through the selected authority path.\n\nMake the action narrow enough to review\n\nA useful packet names the actor and session, project and environment, connector or adapter, target resource, exact operation, bounded parameters, data class, risk, policy profile, approval requirement, expiry, nonce, idempotency key, and rollback or receipt obligation. The policy engine can evaluate those fields, but the enforcement layer must reject substitutions and missing evidence.\n\nA human approval should bind the same digest and scope that the authority decision will consume. If a target, payload, resource, capability, or policy version changes, the system creates a new decision rather than silently reusing the old answer. This is where fail-closed behavior becomes operational instead of rhetorical.\n\nTest the integration as a chain\n\nTest policy inputs with allow, deny, approval-required, unknown-capability, expired, and malformed cases. Test identity mismatch, stale session, wrong audience, resource substitution, scope widening, duplicate idempotency key, connector timeout, partial external result, and revoked authorization. Assert that no adapter runs without a matching authorization bundle.\n\nThen inspect the receipt: requested, approved, and executed digests should agree; adapter and sandbox identity should be present; timing and bounded result should be recorded; unknown outcome should remain unresolved until reconciliation. OPA, Cedar, OpenFGA, Entra ID, Okta, Auth0, and Keycloak can each contribute useful signals. Rangoon’s launch architecture keeps them behind the same exact-action boundary.\n\nSources\nOpen Policy Agent documentation: https://www.openpolicyagent.org/docs/latest/\nCedar policy language reference: https://docs.cedarpolicy.com/",
      "image": "https://rangoon.ai/assets/editorial/article-governance-social.webp",
      "date_published": "2026-10-03T00:00:00Z",
      "date_modified": "2026-10-03T00:00:00Z",
      "tags": [
        "governance",
        "Guide"
      ]
    }
  ]
}
