Repository navigation
Conversation
…gentcore cli for ACR deployment; fix response_schema and chat_template for gpt-oss models in Rollout Gateway.
581f81f to
46b0c28
Compare
|
|
||
| weights = [0.0] * (prompt_length - 1) + [float(m) for m in record.loss_mask] | ||
| logprobs = [0.0] * (prompt_length - 1) + list(record.logprobs) | ||
| advantages = [0.0] * (prompt_length - 1) + [float(advantage)] * response_length |
There was a problem hiding this comment.
Just a clarification question: do all three supported losses honor weights as a loss mask?
Tool/context tokens receive nonzero advantages here; if weights is ignored, their advantages need to be zeroed using record.loss_mask.
There was a problem hiding this comment.
Good point. I thought SM training client would multiply weights and advantage. But better to double guard it.
| ## 2. Prepare training dataset | ||
| Scripts for preparing AgentCore-compatible dataset are in `src/agentcore_rl_toolkit/backends/experimental/sagemaker/prepare_datasets`. For example, to prepare a `gsm8k` dataset, run: | ||
| ```bash | ||
| python src/agentcore_rl_toolkit/backends/experimental/verl/examples/math_agent/preprocess_gsm8k.py \ |
There was a problem hiding this comment.
The gateway-based verl backend graduated in #116, so this script path no longer exists. can use the new sagemaker preparation script?
would be great if you can clean up other stale references to experimental verl too.
There was a problem hiding this comment.
Thanks for the careful review. Removed the stale references
| [verl](/agentcore-rl-toolkit/guides/verl-backend-setup/) backends, there is | ||
| **no GPU cluster required for the training**: SageMaker hosts the policy weights, | ||
| the sampler, and the optimizer behind an SDK, and the RL loop itself runs as a | ||
| plain Python process on your laptop or a small EC2 box. |
There was a problem hiding this comment.
Could we document the current networking requirement here? The ACR runtime needs VPC mode in the same VPC as the training-loop host, with access to the gateway port. Running directly from a laptop isn't supported by this setup?
Issue #, if available:
Description of changes:
What
Adds an experimental SageMaker Training Sessions backend: a GRPO loop that trains an AgentCore Runtime (ACR)-deployed agent with no local GPU cluster. SageMaker hosts the policy weights, the sampler, and the optimizer behind an SDK, so the RL loop itself is a plain single-process asyncio program that can run on a laptop or a small EC2 box. Token capture reuses the in-repo rollout gateway, exactly as the experimental verl backend does.
How it fits together
The only new engine seam is
SageMakerSdkBackend(rollout_gateway/sampling_backends/sagemaker_sdk.py) — a gatewaySamplingBackendover the SageMakerSamplingClientthat mapstoken_ids -> token_ids + logprobs. LikeTinkerSdkBackend, it does not render: the gateway owns tokenization, which keeps loss-masking well-defined and matches the existing placement rule (independently reachable hosted SDK →sampling_backends/).The training loop
backends/experimental/sagemaker/is driven by a single YAML config (config.py/config.yaml.example):train_grpo.pyforward_backward→optim_step→ rebind samplerrollout.pydatum.pyTraceRecord+ advantage → SageMaker training datumconfig.py/config.yaml.exampleprepare_datasets/prepare_gsm8k.pyTests
Tested on GSM8k. Stable training for 200+ steps. Test accuracy increases from ~70% -> ~90% with
gpt-oss-20b.Docs
New
docs/site/src/content/docs/guides/sagemaker-backend-setup.md(installation, config reference, what the loop does step by step, current limits), wired into the sidebar and linked fromindex.mdxandguides/overview.mdxalongside slime / rllm / verl. A shortSETUP.mdsits next to the code for readers who get there from the source tree.By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.