A GitHub Actions job that mints its own Cognito M2M token to test what it just deployed

GitHub OIDC gives CI an AWS identity, but a JWT-authorized MCP endpoint needs a second one: a Cognito client_credentials token read from terraform output.

The Model Context Protocol M-shaped mark and the Model Context Protocol wordmark in black on a white card
On this page

Introduction

I was asked to get a Model Context Protocol server deployed and verified from CI. Deploying it was one problem; calling it afterwards turned out to be a completely separate one, because the credential that let CI create the endpoint could not be used to talk to it.

That is not a quirk of one product. A pipeline like this performs two different acts: it changes infrastructure, which the cloud provider authorizes, and it uses an application, which the application authorizes. Those are two authorities, so they need two credentials.

This article covers both mechanisms in order: GitHub OIDC for the AWS control plane, then a Cognito client_credentials token for the endpoint. It ends with the mistake that made the resulting test report success while checking nothing.

Why one credential is not enough

Amazon Bedrock AgentCore Runtime, which hosts the server, accepts only OAuth/JWT for inbound calls. There is no SigV4 path, so AWS credentials cannot invoke it at all. An unauthenticated request comes back as 401 with a WWW-Authenticate: Bearer header.

So the AWS identity CI already had could create the runtime, read its ARN, and change its configuration, and could not make a single call to it.

Two identities in one CI jobA GitHub Actions job card on the left holds two separate credentials. The upper one, AWS credentials obtained by exchanging an OIDC token, connects to an AWS control plane card. The lower one, a Cognito machine-to-machine access token, connects to an AgentCore MCP endpoint card. The endpoint card carries a gate strip saying it accepts no SigV4 path and that OAuth/JWT is its only inbound authentication, so the AWS credential cannot reach it.GitHub Actions jobOIDC token exchangedfor AWS credentialsCognito M2Maccess tokenAssumeRoleWithWebIdentity,signed with SigV4AWS control planeterraform apply, read outputsAuthorization: Bearer header,checked against allowedClientsAgentCore MCP endpointBearer token on every callNo SigV4 path here.OAuth/JWT is the only inbound auth.
One job, two credentials. The AWS identity reaches the control plane; only the OAuth identity reaches the endpoint.

Worth being precise: this is not two AWS roles. One credential is an AWS principal. The other is an OAuth access token issued by a user pool, which AWS IAM plays no part in validating.

Mechanism 1: GitHub OIDC, for the AWS control plane

The problem OIDC solves is that CI needs AWS credentials and should not store any. Instead, GitHub issues the running job a short-lived JSON Web Token describing itself, and AWS trades that token for temporary credentials through sts:AssumeRoleWithWebIdentity.

Two things have to exist in the AWS account. An IAM OIDC identity provider for token.actions.githubusercontent.com, created once. And a role whose trust policy accepts tokens from it:

trust-policy.json
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": "repo:my-org/my-repo:environment:production"
}
}
}]
}

How you create those two resources does not matter much, and is where a lot of write-ups get specific for no reason. What matters is the sub condition on the last line.

The sub claim describes where the job ran

sub is GitHub’s statement about the workload asking for credentials, and it takes a different form depending on how the job was triggered:

How the job ransub value
References a GitHub Environmentrepo:ORG/REPO:environment:production
Triggered by a pull request event, no environmentrepo:ORG/REPO:pull_request
Neither of the aboverepo:ORG/REPO:ref:refs/heads/main

Environment wins over the others when present. A pull request run produces the same pull_request value for every PR, with no number in it, because the per-run detail lives in the ref and job_workflow_ref claims instead.

The three forms of the OIDC subject claimThree rows, one per trigger context. A job referencing a GitHub Environment produces a subject ending in colon environment colon production, which is accepted. A pull request event without an Environment produces one ending in colon pull_request, which is rejected. Anything else produces one ending in colon ref colon refs slash heads slash main, also rejected. A note explains that widening the condition with a wildcard would accept all three, including runs from forks.One workflow, three possible subjectsreferences a GitHubEnvironmentrepo:ORG/REPO:environment:productionacceptedpull request event,no Environmentrepo:ORG/REPO:pull_requestrejectedneither of the aboverepo:ORG/REPO:ref:refs/heads/mainrejectedA trust condition pinned to the first form accepts that one string and nothing else.Widening it with a wildcard would accept all three, including runs from forks.
The same workflow produces a different subject depending on its trigger. A trust condition pinned to one form rejects the other two.

The practical consequence is that a trust condition and a workflow trigger have to be designed together. Pin the condition to environment:production and the job must declare that environment, or sub comes out as a ref and the role refuses it. The failure arrives as a generic access-denied at assume-role time, which says nothing about which claim mismatched.

Pinning to an environment is the recommended choice rather than merely one option: a GitHub Environment can carry required reviewers and branch restrictions, so the approval gate and the trust condition describe the same thing.

Security reviews flag that pattern specifically, along with omitting the sub condition altogether, which leaves the role assumable by any workflow on GitHub. If one role genuinely has to serve several contexts, list them with StringLike over values you have chosen rather than over *.

Repositories created after 15 July 2026 also get an immutable default subject containing the owner and repository IDs rather than their names, so a rename no longer breaks the trust policy. Worth knowing which format your repository uses before writing the condition.

The workflow side is small. Two lines carry it:

.github/workflows/deploy.yml
jobs:
deploy:
environment:
name: production # decides the sub claim's third segment
permissions:
id-token: write # without this, no OIDC token is issued at all
steps:
- uses: aws-actions/configure-aws-credentials@<pinned-sha>
with:
role-to-assume: arn:aws:iam::111122223333:role/ci-deploy
aws-region: ap-northeast-1

id-token: write is the permission to request a token, not to write anything. Omitting it is the other quiet failure here, and it looks identical to a bad trust policy from the job’s side.

Mechanism 2: a client_credentials token, for the endpoint

Now the second authority. The endpoint validates a JWT from a Cognito user pool, so CI needs a token from that pool.

The obvious approach is the wrong one. A user pool’s password grant authenticates a person, so using it from CI means creating a user, giving it a static password, and storing that password where the pipeline can read it. It also puts a password into infrastructure state if you create the user in code. Every one of those costs is permanent.

client_credentials exists for exactly this case: it authenticates an application rather than a person, so there is no user in the picture at all. The grant needs a resource server to own the scope being requested, and a client that is allowed to request it:

cognito.tf
resource "aws_cognito_resource_server" "api" {
identifier = "mcp-hub"
user_pool_id = aws_cognito_user_pool.this.id
scope {
scope_name = "invoke"
scope_description = "Invoke the runtime"
}
}
resource "aws_cognito_user_pool_client" "machine" {
name = "ci"
user_pool_id = aws_cognito_user_pool.this.id
generate_secret = true
allowed_oauth_flows = ["client_credentials"]
allowed_oauth_flows_user_pool_client = true
allowed_oauth_scopes = ["${aws_cognito_resource_server.api.identifier}/invoke"]
access_token_validity = 1
token_validity_units {
access_token = "hours"
}
}

Two details worth knowing before you write this. openid, profile and email are not valid scopes for client_credentials, so only a resource server’s custom scope belongs in that list. And the flow issues no refresh token, so there is nothing to configure for one.

The same shape as STS

This is the part that makes the two mechanisms feel like one idea rather than two chores.

Both exchange a durable identity for a credential that expires. AssumeRoleWithWebIdentity takes GitHub’s token and returns AWS credentials that last the job. The token endpoint takes a client ID and secret and returns an access token that lasts an hour. In both cases the thing presented on the wire is short-lived, and in both cases CI stores no long-term credential of its own.

Terminal window
curl -sS -X POST "https://<domain>.auth.<region>.amazoncognito.com/oauth2/token" \
-H "Content-Type: application/x-www-form-urlencoded" \
-u "${CLIENT_ID}:${CLIENT_SECRET}" \
-d "grant_type=client_credentials" \
--data-urlencode "scope=mcp-hub/invoke"
How CI obtains and presents a machine tokenFive numbered steps running downward. First, CI reads the client id and secret from Terraform state. Second, it posts to the Cognito token endpoint with the client credentials grant and the resource server scope. Third, it receives an access token carrying a client_id claim and no audience claim. Fourth, it calls the runtime with an Authorization Bearer header. Fifth, AgentCore validates the client_id against its allowedClients list. A note at the foot says this is the same trade STS makes: one durable credential held in an encrypted backend, and only short-lived tokens on the wire, so CI stores no credential of its own.1Read the client credentials from stateterraform output -raw m2m_client_id / m2m_client_secret2POST to the Cognito token endpointgrant_type=client_credentials, scope=mcp-hub/invoke3Receive a short-lived access tokencarries a client_id claim, no aud claim, expires in 1 hour4Call the runtime with a Bearer headerAuthorization: Bearer, over plain HTTPS5The authorizer validates the tokenclient_id checked against the allowed listThe same trade STS makes: one durable credential held in an encrypted backend,and only short-lived tokens on the wire. CI stores no credential of its own.
The exchange. After the credentials are read, nothing in this path involves AWS at all.

The client secret is the one durable credential in the design, and it does not have to live in CI. Terraform generates it and holds it in state, so the job that already has state access can read it and hand it to the next step:

Terminal window
secret=$(terraform output -raw m2m_client_secret)
echo "::add-mask::${secret}" # or it lands in the log
echo "CLIENT_SECRET=${secret}" >> "$GITHUB_ENV"
echo "CLIENT_ID=$(terraform output -raw m2m_client_id)" >> "$GITHUB_ENV"

That is worth stating as a trade rather than a win: the secret is now sensitive data at rest in state, which is acceptable when the state backend is encrypted and access-controlled, and is not acceptable if it is a file on someone’s laptop.

The endpoint side is one line of configuration. AgentCore validates the token’s client_id claim against a list, which is what makes a machine token usable here. A client_credentials token carries client_id and carries no aud:

authorizer_configuration {
custom_jwt_authorizer {
discovery_url = "https://cognito-idp.<region>.amazonaws.com/<pool-id>/.well-known/openid-configuration"
allowed_clients = [aws_cognito_user_pool_client.machine.id]
}
}

Making the smoke test tell the truth

With both mechanisms working, the test was: open a session, list the tools, call a trivial one, then call one that reaches an external API.

The agent wired that last call up behind a flag, switched off. The endpoint had no configured egress path at the time, so the call was expected to fail, and skipping it kept the pipeline green. I rejected that, because a run which skips its only real dependency is not a result:

the only thing I care is if it really tell us the mcp works or not

The flag came out. To keep that safe I also asked for the step to report rather than gate, since a red build for an unresolved networking question teaches people to ignore CI:

- name: Invoke the endpoint
continue-on-error: true

Then two bugs in the test itself, both of which produced failures that looked like findings about the system under test.

The token was minted and thrown away. The retry loop read:

Terminal window
BEARER_TOKEN="$(./token.sh)" && python3 client.py

VAR=value cmd prefixes an assignment onto a command and passes it through. VAR=value && cmd is an assignment statement followed by a separate command, so the variable is set in the shell and never exported. bash -n cannot catch it, because the syntax is valid:

Terminal window
IN
BEARER="$(echo tok-123)" && python3 -c 'import os; print(repr(os.getenv("BEARER")))'
OUT
None
Terminal window
IN
if t="$(echo tok-123)"; then export BEARER="$t"; python3 -c 'import os; print(repr(os.getenv("BEARER")))'; fi
OUT
'tok-123'

The error was parsed away. With the token exported, the call ran and died on json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0). When a tool raises inside an MCP server, the exception message comes back as the tool’s content. Parsing that content as JSON destroys the error and replaces it with a complaint about its first character:

payload = result.content[0].text if result.content else ""
if result.isError or not payload.lstrip().startswith("{"):
print(f"tool returned an error: {payload!r}")
sys.exit(1)
data = json.loads(payload)

The run that failed on its last line

The run after both fixes reported something true:

token obtained (858 chars)
tools: 5 listed
ping: ok
call_external_api:
tool returned an error: 'Error executing tool ...: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate (_ssl.c:1032)'

Each line depends on strictly more of the stack than the one above it. That is what makes the first failure informative instead of merely red:

LineWhat passing it rules out
token obtainedThe machine client, the resource server, the scope, and the secret plumbing are all correct
tools: 5 listedThe JWT authorizer accepted the client_id claim, and the protocol handshake completes
ping: okThe image was pulled from wherever it lives, the container started, and it is running our code
the external callNothing: this is the layer that failed

So the failure localised itself. The outbound call reached a TLS handshake and rejected the certificate, which is neither a timeout nor a DNS failure: egress works, and something on the path is substituting the certificate. Three of the four layers were no longer in question, and the fourth came with an error string specific enough to search.

That ordering is worth designing rather than stumbling into. A test whose steps each add one dependency tells you where it broke; one composite call only tells you that it broke. The cost is that the steps have to be genuinely nested, so a step that shares no dependency with the one before it buys nothing.

Summary

Two mechanisms, because there are two authorities. The cloud provider decides who may change infrastructure, and the application decides who may call it. GitHub OIDC covers the first by exchanging a job-describing token for temporary AWS credentials. The sub claim in that token has to match the trust condition you wrote, so the condition and the workflow trigger are one design decision rather than two.

client_credentials covers the second, and its value is what it leaves out. Authenticating an application instead of a person removes the user, the static password, and the place you would have had to keep it. What is left is the same trade STS makes: hold one durable credential somewhere protected, present only short-lived tokens.

The last part generalises further than the rest. Order the checks so each depends on more of the stack than the last, and a failure names the layer it lives in. Ours cleared three layers and stopped on the fourth, which was the only one whose answer nobody knew. With the skip flag in place it would have reported four passes while exercising three.

References

Share this article