-
Notifications
You must be signed in to change notification settings - Fork 10
feat(continuous-deployment): Update deployment scripts #1191
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
64839a4
60185fe
3c7a560
3703ab0
c4f492d
254f6b5
c71f66e
a18f2bf
5a029bb
188629a
e1f53dc
6213186
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,60 @@ | ||
| name: Discord deploy notification | ||
| description: Post a green (healthy) or red (failed) deploy-status embed to Discord. | ||
|
|
||
| inputs: | ||
| webhook-url: | ||
| description: Discord webhook URL. When empty, the step is a no-op. | ||
| required: true | ||
| name: | ||
| description: Deploy target name shown in the title (e.g. kaapi-staging). | ||
| required: true | ||
| release: | ||
| description: Release identifier (tag or run number). | ||
| required: true | ||
| sha: | ||
| description: Full commit SHA; truncated to 7 chars for display. | ||
| required: true | ||
| run-url: | ||
| description: Link to the workflow run. | ||
| required: true | ||
| ok: | ||
| description: "'true' for a healthy deploy; anything else renders as failed." | ||
| required: true | ||
| failure-reason: | ||
| description: Optional extra line explaining a failure (e.g. CI blocked the release). | ||
| required: false | ||
| default: "" | ||
|
|
||
| runs: | ||
| using: composite | ||
| steps: | ||
| - shell: bash | ||
| env: | ||
| DISCORD_WEBHOOK_URL: ${{ inputs.webhook-url }} | ||
| NAME: ${{ inputs.name }} | ||
| RELEASE: ${{ inputs.release }} | ||
| SHA: ${{ inputs.sha }} | ||
| RUN_URL: ${{ inputs.run-url }} | ||
| OK: ${{ inputs.ok }} | ||
| REASON: ${{ inputs.failure-reason }} | ||
| run: | | ||
| [ -z "$DISCORD_WEBHOOK_URL" ] && { echo "No webhook configured, skipping"; exit 0; } | ||
| if [ "$OK" = "true" ]; then | ||
| TITLE="🟢 $NAME deployment healthy"; COLOR=3066993 # green | ||
| else | ||
| TITLE="🔴 $NAME deployment failed"; COLOR=15158332 # red | ||
| fi | ||
| SHA_SHORT=$(echo "$SHA" | cut -c1-7) | ||
|
|
||
| fields=$(jq -n --arg release "$RELEASE" --arg sha "$SHA_SHORT" \ | ||
| '[{name:"Release",value:$release,inline:true},{name:"SHA",value:$sha,inline:true}]') | ||
| # Only a failed deploy with a stated cause gets the extra Reason line. | ||
| if [ "$OK" != "true" ] && [ -n "$REASON" ]; then | ||
| fields=$(echo "$fields" | jq --arg r "$REASON" '. + [{name:"Reason",value:$r,inline:false}]') | ||
| fi | ||
|
|
||
| payload=$(jq -n --arg title "$TITLE" --argjson color "$COLOR" \ | ||
| --arg url "$RUN_URL" --argjson fields "$fields" \ | ||
| '{embeds:[{title:$title,url:$url,color:$color,fields:$fields,timestamp:(now|todate)}]}') | ||
| curl -sf -H "Content-Type: application/json" -X POST -d "$payload" "$DISCORD_WEBHOOK_URL" \ | ||
| || echo "Discord notification failed to send" | ||
| Original file line number | Diff line number | Diff line change | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| @@ -0,0 +1,53 @@ | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| name: Deploy ECS service and verify rollout | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| description: >- | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Force a new deployment on one ECS service (turning on the deployment circuit | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| breaker with rollback) and wait until its rollout COMPLETED or FAILED. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| inputs: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| cluster: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| description: ECS cluster name. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| required: true | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| service: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| description: ECS service name. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| required: true | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| task-definition: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| description: >- | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Task-definition family (no revision). When set, rolls the service to the | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| family's latest active revision; when empty, keeps the service's current | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| task definition. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| required: false | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| default: "" | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| poll-interval: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| description: Seconds between rollout-state polls. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| required: false | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| default: "15" | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| runs: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| using: composite | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| steps: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| - shell: bash | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| env: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| CLUSTER: ${{ inputs.cluster }} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| SERVICE: ${{ inputs.service }} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| FAMILY: ${{ inputs.task-definition }} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| POLL_INTERVAL: ${{ inputs.poll-interval }} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| run: | | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| args=(--cluster "$CLUSTER" --service "$SERVICE" --force-new-deployment | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| --deployment-configuration '{"deploymentCircuitBreaker":{"enable":true,"rollback":true},"maximumPercent":200,"minimumHealthyPercent":100}') | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. here, so we have enabled the deployment circuit breaker, so if any new task fails during deployment, it will automatically roll back to the previous healthy version, ensuring that the service remains available.
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. so in this configuration, I have set I have set so, ECS will only stop the old tasks after the new tasks are healthy, ensuring zero downtime and that the service never drops below its required capacity. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| [ -n "$FAMILY" ] && args+=(--task-definition "$FAMILY") | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| echo "[$SERVICE] forcing new deployment" | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| aws ecs update-service "${args[@]}" >/dev/null | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| while true; do | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| STATE=$(aws ecs describe-services --cluster "$CLUSTER" --services "$SERVICE" \ | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| --query "services[0].deployments[?status=='PRIMARY'].rolloutState | [0]" \ | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| --output text) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| case "$STATE" in | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| COMPLETED) echo "[$SERVICE] rollout COMPLETED"; break ;; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| FAILED) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| echo "::error::[$SERVICE] rollout FAILED — new tasks never became healthy (rolled back by circuit breaker)" | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| exit 1 ;; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| *) echo "[$SERVICE] rollout $STATE — waiting ${POLL_INTERVAL}s"; sleep "$POLL_INTERVAL" ;; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| esac | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| done | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Comment on lines
+40
to
+53
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift 🔎 Supported by static analysis🏁 Script executed: nl -ba .github/actions/ecs-deploy/action.yml | sed -n '1,100p'
nl -ba .github/workflows/create-release.yml | sed -n '100,140p'
nl -ba .github/workflows/deploy-staging-ecs.yml | sed -n '68,90p'Repository: ProjectTech4DevAI/kaapi-backend Length of output: 5021 🏁 Script executed: set -e
printf '%s\n' '--- changed paths ---'
git diff --stat 361e80dee14487f8b2e1245d2b7ba4c21893adc3 6213186d5bef775ae1cee3379e993e75cb763d50
printf '%s\n' '--- action at base ---'
git show 361e80dee14487f8b2e1245d2b7ba4c21893adc3:.github/actions/ecs-deploy/action.yml 2>/dev/null || true
printf '%s\n' '--- ECS-related callers and inline workflows ---'
rg -n -C 4 'ecs-deploy|update-service|rolloutState|deploymentCircuitBreaker|describe-services|publish' .github
printf '%s\n' '--- relevant workflow diff ---'
git diff --unified=30 361e80dee14487f8b2e1245d2b7ba4c21893adc3 6213186d5bef775ae1cee3379e993e75cb763d50 -- .githubRepository: ProjectTech4DevAI/kaapi-backend Length of output: 25593 🌐 Web query:
💡 Result: 🌐 Web query:
💡 Result: 🌐 Web query:
💡 Result: Poll the deployment returned by
🐛 Suggested fix- aws ecs update-service "${args[@]}" >/dev/null
+ DEPLOYMENT_ID=$(aws ecs update-service "${args[@]}" \
+ --query "service.deployments[?status=='PRIMARY'].id | [0]" --output text)
+ echo "[$SERVICE] deployment $DEPLOYMENT_ID started"
while true; do
STATE=$(aws ecs describe-services --cluster "$CLUSTER" --services "$SERVICE" \
- --query "services[0].deployments[?status=='PRIMARY'].rolloutState | [0]" \
+ --query "services[0].deployments[?id=='$DEPLOYMENT_ID'].rolloutState | [0]" \
--output text)
case "$STATE" in
COMPLETED) echo "[$SERVICE] rollout COMPLETED"; break ;;
FAILED)
echo "::error::[$SERVICE] rollout FAILED — new tasks never became healthy (rolled back by circuit breaker)"
exit 1 ;;
+ None) echo "[$SERVICE] deployment state is not visible — waiting ${POLL_INTERVAL}s"; sleep "$POLL_INTERVAL" ;;
*) echo "[$SERVICE] rollout $STATE — waiting ${POLL_INTERVAL}s"; sleep "$POLL_INTERVAL" ;;
esac
doneThe missing-state branch should fail after a bounded retry window or after checking deployment history, rather than failing on the first eventually consistent read. 📝 Committable suggestion
Suggested change
🤖 Prompt for AI Agents |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -46,9 +46,9 @@ jobs: | |
| uses: actions/checkout@v7 | ||
|
|
||
| - name: Configure AWS credentials | ||
| uses: aws-actions/configure-aws-credentials@v6 # More information on this action can be found below in the 'AWS Credentials' section | ||
| uses: aws-actions/configure-aws-credentials@v6 | ||
| with: | ||
| role-to-assume: arn:aws:iam::024209611402:role/github-action-role | ||
| role-to-assume: ${{ secrets.AWS_DEPLOY_ROLE_ARN }} | ||
|
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. when this script was initially written, this value was hardcoded, which isn’t ideal. It should be picked from secrets instead, so if it changes in the future, we can update it easily without having to make changes to the workflow every time. |
||
| aws-region: ap-south-1 | ||
|
|
||
| - name: Login to Amazon ECR | ||
|
|
@@ -62,6 +62,7 @@ jobs: | |
| TAG: ${{ github.ref_name }} | ||
| run: | | ||
| docker build \ | ||
| --build-arg GIT_SHA=${{ github.sha }} \ | ||
| -t $REGISTRY/$REPOSITORY:latest \ | ||
| -t $REGISTRY/$REPOSITORY:$TAG \ | ||
| ./backend | ||
|
|
@@ -103,16 +104,35 @@ jobs: | |
| fi | ||
| echo "Migration completed successfully" | ||
|
|
||
| - name: Deploy to ECS | ||
| run: | | ||
| aws ecs update-service \ | ||
| --cluster ${{ vars.AWS_RESOURCE_PREFIX }}-cluster \ | ||
| --service ${{ vars.AWS_RESOURCE_PREFIX }}-service \ | ||
| --task-definition ${{ vars.AWS_RESOURCE_PREFIX }}-task \ | ||
| --force-new-deployment | ||
|
|
||
| aws ecs update-service \ | ||
| --cluster ${{ vars.AWS_RESOURCE_PREFIX }}-cluster \ | ||
| --service ${{ vars.AWS_RESOURCE_PREFIX }}-celery-task \ | ||
| --task-definition ${{ vars.AWS_RESOURCE_PREFIX }}-celery-task \ | ||
| --force-new-deployment | ||
| - name: Deploy backend service | ||
| timeout-minutes: 15 | ||
| uses: ./.github/actions/ecs-deploy | ||
| with: | ||
| cluster: ${{ vars.AWS_RESOURCE_PREFIX }}-cluster | ||
| service: ${{ vars.AWS_RESOURCE_PREFIX }}-service | ||
| task-definition: ${{ vars.AWS_RESOURCE_PREFIX }}-task | ||
|
|
||
| - name: Deploy celery service | ||
| timeout-minutes: 15 | ||
| uses: ./.github/actions/ecs-deploy | ||
| with: | ||
| cluster: ${{ vars.AWS_RESOURCE_PREFIX }}-cluster | ||
| service: ${{ vars.AWS_RESOURCE_PREFIX }}-celery-task | ||
| task-definition: ${{ vars.AWS_RESOURCE_PREFIX }}-celery-task | ||
|
|
||
| notify: | ||
| needs: [verify-ci, build] | ||
| if: always() | ||
| runs-on: ubuntu-latest | ||
| steps: | ||
| - uses: actions/checkout@v7 | ||
| - uses: ./.github/actions/discord-notify | ||
| with: | ||
| webhook-url: ${{ secrets.DISCORD_WEBHOOK_URL }} | ||
|
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. need to set this discord webhook url in the GitHub action secrets.. |
||
| name: kaapi-production | ||
| release: ${{ github.ref_name }} | ||
| sha: ${{ github.sha }} | ||
| run-url: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} | ||
| ok: ${{ needs.build.result == 'success' }} | ||
| # Call out the specific case where the release was blocked by CI. | ||
| failure-reason: ${{ needs.verify-ci.result != 'success' && 'CI never passed on the tagged commit — release blocked.' || '' }} | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,12 @@ | ||
| from fastapi.testclient import TestClient | ||
|
|
||
| from app.core.config import settings | ||
|
|
||
|
|
||
| def test_health_reports_status_and_sha(client: TestClient) -> None: | ||
| response = client.get("/health") | ||
|
|
||
| assert response.status_code == 200 | ||
| body = response.json() | ||
| assert body["status"] == "ok" | ||
| assert body["sha"] == settings.GIT_SHA |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
Retry transient Discord failures and report exhausted delivery attempts.
If Discord returns HTTP 429,
curl -sffails without retrying. The followingechomakes the action succeed, although Discord received no deployment notification. Add bounded retries that respectRetry-After. If delivery still fails, make the failure visible to the caller rather than reporting a successful action. Discord documents HTTP 429 rate limits, and curl supports retries for that response. (support-dev.discord.com)🤖 Prompt for AI Agents