Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
119 commits
Select commit Hold shift + click to select a range
1e7598c
Fix TP=1 tensor-parallel gather autograd view (#6657)
fwerkor Sep 3, 2026
1030379
Add per-module MFSDP prefetch budgets (#6905)
wujingyue Sep 3, 2026
da533c0
Add flag to disable KV cache scale clamp to ModelOpt export [OMNIML-5…
jenchen13 Sep 3, 2026
ae4d08b
Add HybridModel architecture guidance (#6938)
Phlip79 Sep 3, 2026
b48ab09
Add a payload offload mode with a staging protocol
tdene Sep 2, 2026
789f746
Accept `required_prefix_token_ids` on the endpoint
tdene Sep 1, 2026
a055f8f
Restore the batch schema bridge for hybrid CP sub-samples (#6820)
ilml Sep 3, 2026
7db926d
Add MTP-only training mode (#7021)
YanivDorGalron Sep 3, 2026
4ec08fe
Move MTP inference code to separate mixin class (#7056)
santhnm2 Sep 3, 2026
91149c0
Add a MFSDP v2 port of the deepseek_proxy_fsdp_ep2_fsdp2 functional t…
wujingyue Sep 3, 2026
52500b6
build: AUT-2212 move torch-memory-saver to inference extra (#7068)
svcnemo-autobot Sep 4, 2026
b1233de
perf(refit): reduce MXFP8 conversion overhead (#6952)
wdykas Sep 4, 2026
8da558d
Add current async scheduling pairwise coverage and fixes (#7053)
lmcafee-nvidia Sep 4, 2026
13c5581
test: clarify DeepSeek proxy MFSDP v1 names (#7073)
wujingyue Sep 4, 2026
95387a1
Add support for fused Q Up-Proj GEMM/RoPE/Quant (#6213)
chaseblock Sep 4, 2026
6572312
test(refit): cover DSA models and enable M2N CI (#6998)
wdykas Sep 4, 2026
c171f66
Golden deepseek_proxy_mfsdp_v2_ep2 against its own output; enable det…
wujingyue Sep 4, 2026
1bf32f7
ci: Enhance CI env to check functional tests label (#7072)
balasaajay Sep 4, 2026
9d790cf
Validate num_attention_heads is divisible by num_query_groups in Tran…
huthvincent Sep 4, 2026
fce0676
chore(checkpointing): remove dead load_biencoder_checkpoint and _add_…
Anai-Guo Sep 4, 2026
f8d6302
fix(test): AUT-2220 preserve golden value precision (#7087)
svcnemo-autobot Sep 4, 2026
df4996f
Fix runtime CP group handling in attention for hybrid CP (#6821)
ilml Sep 5, 2026
def07af
fix: restore each rank's own RNG state on checkpoint load (#6858)
ZhiyuLi-Nvidia Sep 5, 2026
3b5556e
Add CuTeDSL GDP chunkwise context parallelism (#7038)
Mellonta Sep 5, 2026
b1fe759
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Sep 5, 2026
b3393bb
Expose MFSDP v2 custom-schedule lifecycle (#6484)
shjwudp Sep 7, 2026
c9b53d0
Fix FSDP mixed-precision reset for fixed buffers (#6924)
jeffnvidia Sep 8, 2026
c768946
add raise_on_error argument (#7110)
dimapihtar Sep 8, 2026
907609b
Add native DeepSeek-V4 hybrid attention orchestration (#6402)
FDecaYed Sep 8, 2026
62fa723
fix: AUT-2228 route partial reruns to the current attempt (#7142)
svcnemo-autobot Sep 8, 2026
154f665
Use metadata dp_cp_group for the tied-embeddings replica_id (#6888)
going-song Sep 8, 2026
7fa83b5
Keep RerunDataIterator wrapping on the hybrid CP data iterator (#6822)
ilml Sep 9, 2026
be85fc5
ci: AUT-2245 set CUDA major for lockfile updates (#7151)
svcnemo-autobot Sep 9, 2026
8dc7ed3
fix(MoE): Preserve exact expert token counts for expert-bias updates …
JF-D Sep 9, 2026
c6be919
chore: rotate oncall schedule
github-actions[bot] Sep 9, 2026
8dcf188
ci: AUT-2250 retry transient coverage artifact uploads (#7159)
svcnemo-autobot Sep 9, 2026
0d2dc9c
Enter both autocast and fp8 contexts in combined 1F1B execution (#5761)
huthvincent Sep 9, 2026
b9c22eb
Derive MFSDP v2 checkpoint chunk metadata from the parameter layout (…
ahmadki Sep 9, 2026
b92c3ed
Add Slack reminders for re-requested reviews (#7071)
Phlip79 Sep 9, 2026
9ca681c
Propagate rope_scaling_factor from args to GPT model construction (#6…
Alvorecer721 Sep 9, 2026
56ef6ed
ci(pr-review): extract review prompts into a non-activating skill (#7…
ko3n1g Sep 9, 2026
32c27fa
feat(inference): add canonical token capture hooks
lauradang Sep 9, 2026
086b755
fix(inference): define prefix replacement boundary helper
lauradang Sep 9, 2026
fc3b3bd
fix(inference): sort capture hook imports
lauradang Sep 9, 2026
5590c4b
First pass at comprehensive user guide (#5325)
shanmugamr1992 Sep 9, 2026
74dc1b5
Support FP4 in 1F1B A2A overlap (#6135)
dingqingy-nv Sep 9, 2026
b2cea22
RL: prevent unnecessary entry into inference mode (#4125)
tdene Sep 9, 2026
ba638f0
Document signing rewritten commits (#7037)
wujingyue Sep 9, 2026
aa9b579
Test SFTTokenizer token loss masking (#7023)
jenchen13 Sep 10, 2026
d455089
Refactor MFSDP parameter group initialization (#7175)
wujingyue Sep 10, 2026
9be4a57
Assert the runtime-CP group contract both ways (#6825)
ilml Sep 10, 2026
eb245dd
fix(inference): serialize merged finished requests
lauradang Sep 10, 2026
cd19bd6
fix(inference): satisfy prefix capture control flow lint
lauradang Sep 10, 2026
8190837
Support full-iteration CUDA graphs with MFSDP v2 (#7075)
wujingyue Sep 10, 2026
b005bf1
fix(gtp): correct expert gradient scaling for GTP/EGTP mismatch (non-…
fanshiqing Sep 10, 2026
c9e7ea7
Fix MFSDP v2 precision-aware optimizer gradient clipping (#7074)
wujingyue Sep 10, 2026
a7ba373
fix(ci): AUT-2260 mirror DCO status before merge queue (#7210)
svcnemo-autobot Sep 10, 2026
abccf5b
Re-derive quantized params from FP32 main params on checkpoint load (…
Connor-XY Sep 10, 2026
fa751da
fix(inference): forward opaque request metadata
lauradang Sep 10, 2026
db79fa5
fix(inference): preserve payload response contracts
lauradang Sep 10, 2026
79be654
Fix order local refit copies after source updates (#6823)
Sunt-ing Sep 10, 2026
8962453
fix(moe): keep precision-overridden grouped experts off the op-fuser …
fanshiqing Sep 10, 2026
c238266
Select kernel backends through BackendSpecProvider, not HAVE_* flags …
guihong-nv Sep 10, 2026
7613fd7
Enable NCCL EP tests (#7195)
YangFei1990 Sep 10, 2026
5bbf727
Fix Transformer Engine API documentation link (#7225)
wujingyue Sep 11, 2026
3a179c1
Restore FusedAdam empty-shard workarounds for TE < 2.18 (#7216)
wujingyue Sep 11, 2026
dd468df
fix: Do not pass is_first_microbatch to TE when it is not maintained …
ZhiyuLi-Nvidia Sep 11, 2026
5762bac
Deduplicate MFSDP context post-backward hooks (#7196)
wujingyue Sep 11, 2026
46fb90c
Raise NotImplementedError from InferenceInterface.base_generate (#7124)
YeonwooSung Sep 11, 2026
2371e90
chore: Update base image to 26.08 (#6991)
balasaajay Sep 11, 2026
5e5325d
Support independent HSDP instances for expert parameters* (#7012)
Autumn1998 Sep 11, 2026
1e270ad
Add attn_logit_softcapping config, plumbed to TE and local attention …
nvegesna-netizen Sep 11, 2026
4605e3f
Fix QAD checkpoint and dataset loading (#7137)
jenchen13 Sep 11, 2026
eb5b23a
Run mFSDP v1 functional tests nightly (#7244)
wujingyue Sep 11, 2026
5ba1941
fix: AUT-2254 support mHC in experimental MoE transformer layers (#7235)
svcnemo-autobot Sep 11, 2026
f6c33bd
test(determinism): add kernel replay tests and a PR coverage gate (#7…
Connor-XY Sep 11, 2026
976f723
Allow context parallelism with MFSDP v2 (#7220)
wujingyue Sep 12, 2026
4fe0daf
Re-enable prefill CUDA graphs (#7243)
santhnm2 Sep 12, 2026
845d9be
Mark non-final microbatches for MFSDP v2 hybrid data parallelism (#7186)
wujingyue Sep 12, 2026
3703d4e
test(inference): cover long coordinator replies (#7240)
svcnemo-autobot Sep 12, 2026
fa618b2
Support FP4 tensors in the MoE paged activation stash (#7026)
michal2409 Sep 14, 2026
88cf159
Add block-atomic DBuffer placement (#6978)
wujingyue Sep 14, 2026
312f867
fix(muon): split the fully-gathered QKV tensor before orthogonalizing…
fanshiqing Sep 14, 2026
2f582e6
Separate DBuffer layout construction from allocation (#7257)
wujingyue Sep 14, 2026
552da39
test: enforce full-precision deterministic functional goldens (#7217)
Connor-XY Sep 14, 2026
e812813
Fix token count dtype for fused MoE auxiliary loss which requires int…
zhongbozhu Sep 14, 2026
325c993
fix(resharding): bound NCCL refit P2P groups to avoid kernel-plan spl…
wdykas Sep 14, 2026
b212e6d
Handle NVRx installs without package version (#7247)
tylerpoon Sep 14, 2026
4e45f55
Fix FlashInfer NVLS routing buffer race (#7263)
wdykas Sep 14, 2026
904ad47
Add DeepSeek-V3 NVFP4 training support (#6841)
denys-fridman Sep 14, 2026
c4997e3
Add multimodal vision token utilities (#7249)
tylerpoon Sep 14, 2026
54c62df
Fused GDN attention support (#6645)
ksivaman Sep 14, 2026
26e82ea
feat(inference): forward the prefix splice boundary to the engine pro…
lauradang Sep 14, 2026
3972b96
Harden dynamic inference request lifecycle handling (#7256)
lmcafee-nvidia Sep 14, 2026
4cdad25
Merge branch 'main' into tde/ledger_capture
lauradang Sep 14, 2026
e7f245e
Guard einops import in multimodal utils
lauradang Sep 14, 2026
0e48cdf
fix(determinism): Separate the MoE aux-loss fusion flag so --determin…
ZhiyuLi-Nvidia Sep 14, 2026
78d917e
fix: make GTP weight rematerialization work end-to-end with tied embe…
shanmugamr1992 Sep 14, 2026
1e137dd
ci: Revert recently merged changes as they cause CI unit test failure…
balasaajay Sep 15, 2026
556c74d
Multimodal: add tokenizer path (#3466)
faradawn Sep 14, 2026
79097b4
[GTP] Symmetric memory registration to use custom VMM allocator (#6956)
prajwal1210 Sep 15, 2026
65fe89a
Add refit codeowners (#7330)
wdykas Sep 15, 2026
ad29cd0
ci: Update golden value for hybrid_dynamic_inference_tp1_pp1_dp8_583m…
chtruong814 Sep 15, 2026
192b18f
Add Style Guide (#7179)
Phlip79 Sep 15, 2026
8349ade
Fix GTP checkpoint load failing across different GTP degrees (#7293)
fanshiqing Sep 15, 2026
088cd1d
Add implementation for ShortcutMoE (#6959)
jiemingz Sep 15, 2026
42b8a69
Do not modify MLA spec at runtime for down proj fusion (#7304)
janEbert Sep 15, 2026
d046b0d
Fix MFSDP v2 parameter metadata preservation (#7230)
Autumn1998 Sep 15, 2026
ce2c8ab
Merge branch 'main' into tde/ledger_capture
lauradang Sep 15, 2026
0116fab
Rename the opaque request_metadata dict to offload_params
lauradang Sep 17, 2026
f050d85
Let the prompt preparer replace prefix tokens with its own splicer
lauradang Sep 17, 2026
cdaf623
Update unit tests for the main merge of the offload wire format
lauradang Sep 17, 2026
8384f1b
Harden prompt preparation failures in the dynamic engine
lauradang Sep 17, 2026
2e00d27
Carry offload_params in its own SUBMIT_REQUEST frame
lauradang Sep 17, 2026
485bd65
Drop prompt tensors from offloaded replies and tidy the offload wire …
lauradang Sep 17, 2026
01f1247
Fix stale variable name in test_payload_offload_mode
lauradang Sep 17, 2026
2800313
Surface stage metadata on /v1/completions and reserve underscore keys…
lauradang Sep 17, 2026
1310225
fix(inference): type prompt preparation results
lauradang Sep 17, 2026
0aa6cce
feat(inference): hand compact prompt ids and media tensors to the pay…
lauradang Sep 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
2 changes: 2 additions & 0 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,8 @@ megatron/core/transformer/fsdp_dtensor_checkpoint.py @NVIDIA/core-adlr @NVIDIA/c

megatron/core/dist_checkpointing/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/dist-checkpointing

megatron/core/resharding/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/refit

megatron/core/optimizer/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/mcore-optimizer

megatron/core/optimizer/distrib_optimizer.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/dist-optimizer
Expand Down
17 changes: 17 additions & 0 deletions .github/actions/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -293,15 +293,32 @@ runs:
fi

- name: Upload coverage
id: upload-coverage
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
if: ${{ always() && steps.check.outputs.coverage_report != 'none' }}
continue-on-error: true
with:
name: ${{ steps.check.outputs.coverage_report }}
path: |
coverage.xml
.coverage
include-hidden-files: true

- name: Back off after coverage upload failure
if: ${{ always() && steps.upload-coverage.outcome == 'failure' }}
shell: bash
run: sleep 10

- name: Retry coverage upload
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
if: ${{ always() && steps.upload-coverage.outcome == 'failure' }}
with:
name: ${{ steps.check.outputs.coverage_report }}-retry
path: |
coverage.xml
.coverage
include-hidden-files: true

- name: Upload logs
id: upload-logs
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
Expand Down
2 changes: 1 addition & 1 deletion .github/copy-pr-bot.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
enabled: true
auto_sync_draft: false
auto_sync_ready: true
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "Connor-XY", "DanialTaheri", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JF-D", "JRD971000", "Leili", "Mellonta", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "WanZzzzzz", "Wohox", "XiaodaNV", "YangFei1990", "ZhiyuLi-Nvidia", "abhinav-khattar", "adistomar", "ahmadki", "alokpathy", "ananthsub", "anlthms", "aroshanghias-nvd", "ashehper", "asolergi-nv", "athitten", "balasaajay", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "desh2608", "dimapihtar", "dingqingy-nv", "duncanriach", "ehosseiniasl", "erhoo82", "ericharper", "fanshiqing", "faradawn", "fitsumreda", "freewym", "frsun-nvda", "gautham-kollu", "gdengk", "goelarushi", "guihong-nv", "guyueh1", "hexinw-nvidia", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiaji-huang", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kajalj22", "kamran-nvidia", "kevalmorabia97", "kevjshih", "kingformatty", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwangjunahn", "kwyss-nvidia", "lauradang", "layalir", "lhb8125", "liding-nv", "lmcafee-nvidia", "maanug-nv", "macandro96", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "minitu", "mkhona-nvidia", "nanz-nv", "niyunsheng", "ntajbakhsh", "nvcsathe", "nvegesna-netizen", "parthmannan", "philipcmonk", "prajwal1210", "pthombre", "rapatel", "rhewett-nv", "rogerwaleffe", "sajadn", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "sheliang-nv", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sraman-rgb", "sudhakarsingh27", "svcnvidia-nemo-ci", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "vasunvidia", "wanyingw", "wdykas", "wujingyue", "xiaoyao0115", "xuantengh", "xuwchen", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yqwangustc", "yueshen2016", "yuzhongw-nvidia", "zhehuaichen", "zhongbozhu"]
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "Connor-XY", "DanialTaheri", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JF-D", "JRD971000", "Leili", "Mellonta", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "WanZzzzzz", "Wohox", "XiaodaNV", "YangFei1990", "ZhiyuLi-Nvidia", "abhinav-khattar", "adistomar", "ahmadki", "alokpathy", "ananthsub", "anlthms", "aroshanghias-nvd", "ashehper", "asolergi-nv", "athitten", "balasaajay", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "desh2608", "dimapihtar", "dingqingy-nv", "duncanriach", "ehosseiniasl", "erhoo82", "ericharper", "fanshiqing", "faradawn", "fitsumreda", "freewym", "frsun-nvda", "gautham-kollu", "gdengk", "goelarushi", "guihong-nv", "guyueh1", "hexinw-nvidia", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiaji-huang", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kajalj22", "kamran-nvidia", "kevalmorabia97", "kevjshih", "kingformatty", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwangjunahn", "kwyss-nvidia", "lauradang", "layalir", "lhb8125", "liding-nv", "lmcafee-nvidia", "maanug-nv", "macandro96", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "minitu", "mkhona-nvidia", "nanz-nv", "nathan-nvidia", "niyunsheng", "ntajbakhsh", "nvcsathe", "nvegesna-netizen", "parthmannan", "philipcmonk", "prajwal1210", "pthombre", "rapatel", "rhewett-nv", "rogerwaleffe", "sajadn", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "sheliang-nv", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sraman-rgb", "sudhakarsingh27", "svcnvidia-nemo-ci", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "vasunvidia", "wanyingw", "wdykas", "wujingyue", "xiaoyao0115", "xuantengh", "xuwchen", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yqwangustc", "yueshen2016", "yuzhongw-nvidia", "zhehuaichen", "zhongbozhu"]
8 changes: 4 additions & 4 deletions .github/oncall_schedule.json
Original file line number Diff line number Diff line change
@@ -1,8 +1,4 @@
[
{
"user": "YangFei1990",
"date": "2026-09-02"
},
{
"user": "asolergi-nv",
"date": "2026-09-09"
Expand Down Expand Up @@ -46,5 +42,9 @@
{
"user": "YangFei1990",
"date": "2026-11-18"
},
{
"user": "asolergi-nv",
"date": "2026-11-25"
}
]
1 change: 1 addition & 0 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ Linked issue: <!-- e.g. Fixes #1234 / Related to #1234 -->

- [ ] I have added relevant unit tests
- [ ] I have added relevant functional tests
- [ ] If this PR adds or changes a GPU kernel (Triton, `jit_fuser`/`torch.compile`, CUDA extension, TE or external-library dispatch, or a scatter/index accumulation), I have added or updated its bit-exact determinism test and registered it in `tests/unit_tests/determinism/kernels/manifest.py` ([guide](https://github.com/NVIDIA/Megatron-LM/blob/main/docs/developer/determinism/testing.md))
- [ ] I have added proper typing to my code [Typing guidelines](https://docs.python.org/3/library/typing.html)
- [ ] I have added relevant documentation
- [ ] I have run the [autoformatter.sh](https://github.com/NVIDIA/Megatron-LM/blob/main/tools/autoformat.sh) on my PR
Expand Down
238 changes: 238 additions & 0 deletions .github/scripts/dco_gate.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,238 @@
#!/usr/bin/env python3
# Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

"""Mirror the trusted DCO App verdict to a repository-owned check."""

import json
import os
import re
import sys
from datetime import datetime, timezone
from urllib.error import HTTPError, URLError
from urllib.parse import quote
from urllib.request import Request, urlopen

DCO_APP_SLUG = "dco"
DCO_CHECK_NAME = "DCO"
GATE_CHECK_NAME = "DCO gate"
_SHA_PATTERN = re.compile(r"^[0-9a-f]{40}$")


class GateError(RuntimeError):
"""Raised when the event or GitHub response cannot be trusted."""


def _validated_sha(value: object, source: str) -> str:
if not isinstance(value, str) or not _SHA_PATTERN.fullmatch(value):
raise GateError(f"{source} has an invalid head SHA")
return value


def validate_trigger(payload: dict[str, object], requested_sha: str | None = None) -> str:
"""Return the target SHA after validating the check-run or manual trigger."""

if requested_sha:
return _validated_sha(requested_sha, "manual request")

check_run = payload.get("check_run")
if not isinstance(check_run, dict):
raise GateError("trusted DCO check_run payload is missing")

app = check_run.get("app")
app_slug = app.get("slug") if isinstance(app, dict) else None
if check_run.get("name") != DCO_CHECK_NAME or app_slug != DCO_APP_SLUG:
raise GateError("event did not originate from the trusted DCO App check")
if check_run.get("status") != "completed":
raise GateError("DCO check run is not completed")
return _validated_sha(check_run.get("head_sha"), "DCO check run")


def select_latest_dco(check_runs: list[dict[str, object]], head_sha: str) -> dict[str, object]:
"""Select the newest completed DCO App check for the exact head SHA."""

trusted = []
for check_run in check_runs:
app = check_run.get("app")
app_slug = app.get("slug") if isinstance(app, dict) else None
if (
check_run.get("name") == DCO_CHECK_NAME
and app_slug == DCO_APP_SLUG
and check_run.get("head_sha") == head_sha
and check_run.get("status") == "completed"
):
trusted.append(check_run)

if not trusted:
raise GateError("no completed DCO App check exists for the requested SHA")
return max(trusted, key=_check_run_id)


def select_existing_gate(
check_runs: list[dict[str, object]], head_sha: str
) -> dict[str, object] | None:
"""Find the newest externally identified DCO gate for this SHA."""

external_id = _gate_external_id(head_sha)
matching = [
check_run
for check_run in check_runs
if check_run.get("name") == GATE_CHECK_NAME
and check_run.get("external_id") == external_id
and check_run.get("head_sha") == head_sha
]
return max(matching, key=_check_run_id) if matching else None


def gate_payload(
source: dict[str, object], head_sha: str, *, include_head: bool
) -> dict[str, object]:
"""Build a fail-closed create or update request for the mirrored gate."""

source_conclusion = source.get("conclusion")
conclusion = "success" if source_conclusion == "success" else "failure"
source_id = _check_run_id(source)
source_url = source.get("html_url")
source_output = source.get("output", {})
source_summary = source_output.get("summary") if isinstance(source_output, dict) else None

if isinstance(source_url, str) and source_url.startswith("https://"):
source_reference = f"[DCO check run {source_id}]({source_url})"
else:
source_reference = f"DCO check run {source_id}"
summary = f"Mirrored `{source_conclusion}` from trusted {source_reference} for `{head_sha}`."
if isinstance(source_summary, str) and source_summary.strip():
summary += f"\n\nDCO App result: {source_summary.strip()}"

now = datetime.now(timezone.utc).isoformat().replace("+00:00", "Z")
payload: dict[str, object] = {
"name": GATE_CHECK_NAME,
"external_id": _gate_external_id(head_sha),
"status": "completed",
"conclusion": conclusion,
"completed_at": now,
"output": {"title": GATE_CHECK_NAME, "summary": summary[:65000]},
}
if isinstance(source_url, str) and source_url.startswith("https://"):
payload["details_url"] = source_url
if include_head:
payload["head_sha"] = head_sha
return payload


def _check_run_id(check_run: dict[str, object]) -> int:
check_run_id = check_run.get("id")
if not isinstance(check_run_id, int) or check_run_id <= 0:
raise GateError("check run has an invalid ID")
return check_run_id


def _gate_external_id(head_sha: str) -> str:
return f"dco-gate:{head_sha}"


def _request_json(
method: str, url: str, token: str, payload: dict[str, object] | None = None
) -> dict[str, object]:
data = json.dumps(payload).encode() if payload is not None else None
request = Request(
url,
data=data,
method=method,
headers={
"Accept": "application/vnd.github+json",
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
"X-GitHub-Api-Version": "2022-11-28",
},
)
try:
with urlopen(request, timeout=30) as response: # nosec B310 - URL is fixed to GitHub API.
result = json.load(response)
except HTTPError as error:
detail = error.read().decode(errors="replace")[:1000]
raise GateError(f"GitHub API returned {error.code}: {detail}") from error
except (URLError, TimeoutError) as error:
raise GateError(f"GitHub API request failed: {error}") from error
if not isinstance(result, dict):
raise GateError("GitHub API returned a non-object response")
return result


def _list_check_runs(
api_url: str, repository: str, head_sha: str, name: str, token: str
) -> list[dict[str, object]]:
check_runs: list[dict[str, object]] = []
for page in range(1, 11):
url = (
f"{api_url}/repos/{repository}/commits/{head_sha}/check-runs"
f"?check_name={quote(name)}&filter=all&per_page=100&page={page}"
)
response = _request_json("GET", url, token)
batch = response.get("check_runs")
if not isinstance(batch, list) or not all(isinstance(item, dict) for item in batch):
raise GateError("GitHub API returned invalid check-run data")
check_runs.extend(batch)
if len(batch) < 100:
return check_runs
raise GateError("check-run pagination exceeded the safety limit")


def publish_gate(
payload: dict[str, object],
repository: str,
api_url: str,
token: str,
requested_sha: str | None = None,
) -> dict[str, object]:
"""Re-read the current DCO result and publish its repository gate."""

head_sha = validate_trigger(payload, requested_sha)
source_runs = _list_check_runs(api_url, repository, head_sha, DCO_CHECK_NAME, token)
source = select_latest_dco(source_runs, head_sha)
gate_runs = _list_check_runs(api_url, repository, head_sha, GATE_CHECK_NAME, token)
existing_gate = select_existing_gate(gate_runs, head_sha)

if existing_gate is None:
url = f"{api_url}/repos/{repository}/check-runs"
return _request_json("POST", url, token, gate_payload(source, head_sha, include_head=True))

gate_id = _check_run_id(existing_gate)
url = f"{api_url}/repos/{repository}/check-runs/{gate_id}"
return _request_json("PATCH", url, token, gate_payload(source, head_sha, include_head=False))


def main() -> int:
"""Publish the mirrored DCO check for the current workflow event."""

try:
event_path = os.environ["GITHUB_EVENT_PATH"]
repository = os.environ["GITHUB_REPOSITORY"]
token = os.environ["GITHUB_TOKEN"]
api_url = os.environ.get("GITHUB_API_URL", "https://api.github.com").rstrip("/")
requested_sha = os.environ.get("DCO_GATE_SHA") or None
with open(event_path, encoding="utf-8") as event_file:
payload = json.load(event_file)
if not isinstance(payload, dict):
raise GateError("event payload is not an object")
result = publish_gate(payload, repository, api_url, token, requested_sha)
sys.stdout.write(f"Published {GATE_CHECK_NAME} check run {result.get('id')}\n")
except (GateError, KeyError, OSError, json.JSONDecodeError) as error:
sys.stderr.write(f"DCO gate failed: {error}\n")
return 1
return 0


if __name__ == "__main__":
raise SystemExit(main())
Loading