AWS launches SageMaker inference skill for coding agents
AWS has launched an Agent Toolkit skill that generates inspectable SageMaker Python code for inference benchmarks, deployment recommendations and run comparisons. Users must explicitly confirm before a test sends traffic to a live endpoint.
AWS has launched an aws-ai-ml skill that lets MCP-compatible coding agents generate Amazon SageMaker AI code for inference benchmarking and optimization. Through natural-language requests, developers can prepare endpoint benchmarks, evaluate deployment options and compare test runs, while execution remains under their AWS credentials.
The Agent Toolkit skill turns those requests into executable SageMaker Python SDK v3 code that users can inspect, modify and run in their own environments. AWS lists Kiro, Claude Code and Codex among compatible agents. In September, AWS released SageMaker deployment skills for coding agents.
Local setup requires AWS CLI 2.35 or newer and uv. AWS also provides a preconfigured SageMaker Studio JupyterLab image. The company instructs users to create a fresh private space because skills synchronize only to private spaces, and an existing modified space might not receive the bundled version.
Benchmarking a live endpoint
For an existing endpoint, the skill can generate a notebook that creates a synthetic workload and starts a load test. The benchmark workflow reports throughput, latency percentiles, time to first token, inter-token latency and concurrency. AWS says the agent must warn the user and obtain explicit confirmation before running a benchmark because the test sends real traffic to the live endpoint.
The standard path requires an InService SageMaker endpoint hosting a large language model with an OpenAI-compatible chat-completions API, an S3 bucket for results and an IAM execution role that can access both. For custom request and response formats, users can instead supply a Jinja2 payload template and a JMESPath response query.
Users control the test through a workload configuration. They can define token distributions and traffic settings inline or supply a representative request dataset from S3. Parameters include concurrency, request count, request rate, streaming and duration. AWS does not specify the skill’s defaults for those settings or the exact confirmation interface.
Recommendations and comparisons
The skill can also generate code to evaluate models from S3, SageMaker JumpStart or the Hugging Face Hub against candidate instances and configurations. For inference recommendations, the user provides a model and workload, chooses cost, latency or throughput as the objective, and selects no more than three instance types. SageMaker measures and ranks the options, but the user chooses the deployment configuration.
For gated Hugging Face models, AWS says the agent surfaces the model license and asks the user to accept it and provide a Hugging Face token before staging and evaluation. Given two benchmark job names, the skill can generate a comparison of throughput, latency percentiles and time-to-first-token changes. If a named run is missing, AWS says the agent offers to run it first.
AWS’s public materials leave one boundary unclear. The launch post says model deployment is outside this inference-optimization experience and that the agent can generate a deployment configuration instead. The current umbrella skill instructions, however, route requests to deploy, benchmark or optimize a model into a model-deployment workflow.
The generated code uses the user’s AWS credentials, which must permit the SageMaker operations involved, including creating endpoints and running benchmark or recommendation jobs. AWS says the skill adds no separate IAM configuration. The company instructs users to delete endpoints, remove job objects from the default SageMaker S3 bucket, and stop or delete Studio spaces to avoid continuing charges. AWS says inference recommendations carry no additional service fee, although they can provision billable on-demand compute.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

AWS adds instance preference lists to SageMaker AI jobs

AWS maps Claude 5.5 access across Bedrock endpoints in GovCloud
