A GitHub Actions job that mints its own Cognito M2M token to test what it just deployed
GitHub OIDC gives CI an AWS identity, but a JWT-authorized MCP endpoint needs a second one: a Cognito client_credentials token read from terraform output.

On this page
Introduction
I was asked to get a Model Context Protocol server deployed and verified from CI. Deploying it was one problem; calling it afterwards turned out to be a completely separate one, because the credential that let CI create the endpoint could not be used to talk to it.
That is not a quirk of one product. A pipeline like this performs two different acts: it changes infrastructure, which the cloud provider authorizes, and it uses an application, which the application authorizes. Those are two authorities, so they need two credentials.
This article covers both mechanisms in order: GitHub OIDC for the AWS control plane, then a Cognito client_credentials token for the endpoint. It ends with the mistake that made the resulting test report success while checking nothing.
Why one credential is not enough
Amazon Bedrock AgentCore Runtime, which hosts the server, accepts only OAuth/JWT for inbound calls. There is no SigV4 path, so AWS credentials cannot invoke it at all. An unauthenticated request comes back as 401 with a WWW-Authenticate: Bearer header.
So the AWS identity CI already had could create the runtime, read its ARN, and change its configuration, and could not make a single call to it.
Worth being precise: this is not two AWS roles. One credential is an AWS principal. The other is an OAuth access token issued by a user pool, which AWS IAM plays no part in validating.
Mechanism 1: GitHub OIDC, for the AWS control plane
The problem OIDC solves is that CI needs AWS credentials and should not store any. Instead, GitHub issues the running job a short-lived JSON Web Token describing itself, and AWS trades that token for temporary credentials through sts:AssumeRoleWithWebIdentity.
Two things have to exist in the AWS account. An IAM OIDC identity provider for token.actions.githubusercontent.com, created once. And a role whose trust policy accepts tokens from it:
{ "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": { "Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com", "token.actions.githubusercontent.com:sub": "repo:my-org/my-repo:environment:production" } } }]}How you create those two resources does not matter much, and is where a lot of write-ups get specific for no reason. What matters is the sub condition on the last line.
The sub claim describes where the job ran
sub is GitHub’s statement about the workload asking for credentials, and it takes a different form depending on how the job was triggered:
| How the job ran | sub value |
|---|---|
| References a GitHub Environment | repo:ORG/REPO:environment:production |
| Triggered by a pull request event, no environment | repo:ORG/REPO:pull_request |
| Neither of the above | repo:ORG/REPO:ref:refs/heads/main |
Environment wins over the others when present. A pull request run produces the same pull_request value for every PR, with no number in it, because the per-run detail lives in the ref and job_workflow_ref claims instead.
The practical consequence is that a trust condition and a workflow trigger have to be designed together. Pin the condition to environment:production and the job must declare that environment, or sub comes out as a ref and the role refuses it. The failure arrives as a generic access-denied at assume-role time, which says nothing about which claim mismatched.
Pinning to an environment is the recommended choice rather than merely one option: a GitHub Environment can carry required reviewers and branch restrictions, so the approval gate and the trust condition describe the same thing.
Security reviews flag that pattern specifically, along with omitting the sub condition altogether, which leaves the role assumable by any workflow on GitHub. If one role genuinely has to serve several contexts, list them with StringLike over values you have chosen rather than over *.
Repositories created after 15 July 2026 also get an immutable default subject containing the owner and repository IDs rather than their names, so a rename no longer breaks the trust policy. Worth knowing which format your repository uses before writing the condition.
The workflow side is small. Two lines carry it:
jobs: deploy: environment: name: production # decides the sub claim's third segment permissions: id-token: write # without this, no OIDC token is issued at all steps: - uses: aws-actions/configure-aws-credentials@<pinned-sha> with: role-to-assume: arn:aws:iam::111122223333:role/ci-deploy aws-region: ap-northeast-1id-token: write is the permission to request a token, not to write anything. Omitting it is the other quiet failure here, and it looks identical to a bad trust policy from the job’s side.
Mechanism 2: a client_credentials token, for the endpoint
Now the second authority. The endpoint validates a JWT from a Cognito user pool, so CI needs a token from that pool.
The obvious approach is the wrong one. A user pool’s password grant authenticates a person, so using it from CI means creating a user, giving it a static password, and storing that password where the pipeline can read it. It also puts a password into infrastructure state if you create the user in code. Every one of those costs is permanent.
client_credentials exists for exactly this case: it authenticates an application rather than a person, so there is no user in the picture at all. The grant needs a resource server to own the scope being requested, and a client that is allowed to request it:
resource "aws_cognito_resource_server" "api" { identifier = "mcp-hub" user_pool_id = aws_cognito_user_pool.this.id
scope { scope_name = "invoke" scope_description = "Invoke the runtime" }}
resource "aws_cognito_user_pool_client" "machine" { name = "ci" user_pool_id = aws_cognito_user_pool.this.id
generate_secret = true allowed_oauth_flows = ["client_credentials"] allowed_oauth_flows_user_pool_client = true allowed_oauth_scopes = ["${aws_cognito_resource_server.api.identifier}/invoke"]
access_token_validity = 1 token_validity_units { access_token = "hours" }}Two details worth knowing before you write this. openid, profile and email are not valid scopes for client_credentials, so only a resource server’s custom scope belongs in that list. And the flow issues no refresh token, so there is nothing to configure for one.
The same shape as STS
This is the part that makes the two mechanisms feel like one idea rather than two chores.
Both exchange a durable identity for a credential that expires. AssumeRoleWithWebIdentity takes GitHub’s token and returns AWS credentials that last the job. The token endpoint takes a client ID and secret and returns an access token that lasts an hour. In both cases the thing presented on the wire is short-lived, and in both cases CI stores no long-term credential of its own.
curl -sS -X POST "https://<domain>.auth.<region>.amazoncognito.com/oauth2/token" \ -H "Content-Type: application/x-www-form-urlencoded" \ -u "${CLIENT_ID}:${CLIENT_SECRET}" \ -d "grant_type=client_credentials" \ --data-urlencode "scope=mcp-hub/invoke"The client secret is the one durable credential in the design, and it does not have to live in CI. Terraform generates it and holds it in state, so the job that already has state access can read it and hand it to the next step:
secret=$(terraform output -raw m2m_client_secret)echo "::add-mask::${secret}" # or it lands in the logecho "CLIENT_SECRET=${secret}" >> "$GITHUB_ENV"echo "CLIENT_ID=$(terraform output -raw m2m_client_id)" >> "$GITHUB_ENV"That is worth stating as a trade rather than a win: the secret is now sensitive data at rest in state, which is acceptable when the state backend is encrypted and access-controlled, and is not acceptable if it is a file on someone’s laptop.
The endpoint side is one line of configuration. AgentCore validates the token’s client_id claim against a list, which is what makes a machine token usable here. A client_credentials token carries client_id and carries no aud:
authorizer_configuration { custom_jwt_authorizer { discovery_url = "https://cognito-idp.<region>.amazonaws.com/<pool-id>/.well-known/openid-configuration" allowed_clients = [aws_cognito_user_pool_client.machine.id] }}Making the smoke test tell the truth
With both mechanisms working, the test was: open a session, list the tools, call a trivial one, then call one that reaches an external API.
The agent wired that last call up behind a flag, switched off. The endpoint had no configured egress path at the time, so the call was expected to fail, and skipping it kept the pipeline green. I rejected that, because a run which skips its only real dependency is not a result:
the only thing I care is if it really tell us the mcp works or not
The flag came out. To keep that safe I also asked for the step to report rather than gate, since a red build for an unresolved networking question teaches people to ignore CI:
- name: Invoke the endpoint continue-on-error: trueThen two bugs in the test itself, both of which produced failures that looked like findings about the system under test.
The token was minted and thrown away. The retry loop read:
BEARER_TOKEN="$(./token.sh)" && python3 client.pyVAR=value cmd prefixes an assignment onto a command and passes it through. VAR=value && cmd is an assignment statement followed by a separate command, so the variable is set in the shell and never exported. bash -n cannot catch it, because the syntax is valid:
INBEARER="$(echo tok-123)" && python3 -c 'import os; print(repr(os.getenv("BEARER")))'OUTNoneINif t="$(echo tok-123)"; then export BEARER="$t"; python3 -c 'import os; print(repr(os.getenv("BEARER")))'; fiOUT'tok-123'The error was parsed away. With the token exported, the call ran and died on json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0). When a tool raises inside an MCP server, the exception message comes back as the tool’s content. Parsing that content as JSON destroys the error and replaces it with a complaint about its first character:
payload = result.content[0].text if result.content else ""
if result.isError or not payload.lstrip().startswith("{"): print(f"tool returned an error: {payload!r}") sys.exit(1)
data = json.loads(payload)The run that failed on its last line
The run after both fixes reported something true:
token obtained (858 chars)tools: 5 listedping: okcall_external_api: tool returned an error: 'Error executing tool ...: [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate (_ssl.c:1032)'Each line depends on strictly more of the stack than the one above it. That is what makes the first failure informative instead of merely red:
| Line | What passing it rules out |
|---|---|
token obtained | The machine client, the resource server, the scope, and the secret plumbing are all correct |
tools: 5 listed | The JWT authorizer accepted the client_id claim, and the protocol handshake completes |
ping: ok | The image was pulled from wherever it lives, the container started, and it is running our code |
| the external call | Nothing: this is the layer that failed |
So the failure localised itself. The outbound call reached a TLS handshake and rejected the certificate, which is neither a timeout nor a DNS failure: egress works, and something on the path is substituting the certificate. Three of the four layers were no longer in question, and the fourth came with an error string specific enough to search.
That ordering is worth designing rather than stumbling into. A test whose steps each add one dependency tells you where it broke; one composite call only tells you that it broke. The cost is that the steps have to be genuinely nested, so a step that shares no dependency with the one before it buys nothing.
Summary
Two mechanisms, because there are two authorities. The cloud provider decides who may change infrastructure, and the application decides who may call it. GitHub OIDC covers the first by exchanging a job-describing token for temporary AWS credentials. The sub claim in that token has to match the trust condition you wrote, so the condition and the workflow trigger are one design decision rather than two.
client_credentials covers the second, and its value is what it leaves out. Authenticating an application instead of a person removes the user, the static password, and the place you would have had to keep it. What is left is the same trade STS makes: hold one durable credential somewhere protected, present only short-lived tokens.
The last part generalises further than the rest. Order the checks so each depends on more of the stack than the last, and a failure names the layer it lives in. Ours cleared three layers and stopped on the fourth, which was the only one whose answer nobody knew. With the skip flag in place it would have reported four passes while exercising three.
References
- OpenID Connect reference: the sub claim’s forms and which trigger produces each
- Configuring OpenID Connect in Amazon Web Services, including the trust policy shape
- Avoiding mistakes with AWS OIDC integration conditions, on wildcard subjects and forks
- Configure inbound JWT authorizer for AgentCore, on validating allowedClients against client_id
- Amazon Cognito token endpoint and the client_credentials grant





