Skip to content

添加回退校验+迭代循环+[修改input_case 为从json读取] - #82

Closed
chopper0126 wants to merge 5 commits into
Just-it:mainfrom
chopper0126:br_ascend_claudecode
Closed

chopper0126 wants to merge 5 commits into
Just-it:mainfrom
chopper0126:br_ascend_claudecode

Conversation

@chopper0126

@chopper0126 chopper0126 commented Apr 13, 2026 •

Copy link
Copy Markdown
Collaborator

变更说明

[添加回退校验+迭代循环]
[add batch_run_performance.sh]
[修改input_case 为从json读取]

变更类型

  • 新增算子 / Skill
  • 性能优化
  • Bug 修复
  • Benchmark(新增 case / 修改评测逻辑)
  • Agent / 框架改动
  • 文档 / 基础设施

影响范围

  • Triton 侧
  • AscendC 侧
  • 共享(router / benchmark-scheduler)

性能数据(涉及算子生成/优化时必填)

Benchmark 评测(与 BASELINE.md 对比)

指标 BASELINE 本次评测
编译通过数 4 13
精度通过数 4 13
平均 Speedup

测试环境

  • 设备型号:
  • CANN 版本:
  • PyTorch 版本:

冒烟测试(涉及算子生成/框架改动时必填)

  • Triton 通路:✅ / ❌(失败原因:)
  • AscendC 通路:✅ / ❌(失败原因:)

验证清单

  • 双通路冒烟测试通过
  • 通过率不退化(编译、精度均 >= BASELINE)
  • 平均 Speedup 不退化(>= BASELINE × 0.95)
  • 性能优化类:已跑全量评测、逐任务无退化、至少 1 个提升 >= 5%

退化说明(如有通过率下跌)

@chopper0126

Copy link
Copy Markdown
Collaborator Author
>>> 锁定设备 1 并启动算子: 10_SwigluQuant
>>> 锁定设备 2 并启动算子: 13_InterleaveRope
>>> 锁定设备 3 并启动算子: 14_AdaptiveInstanceNormalization2DBackward
>>> 锁定设备 4 并启动算子: 18_FusedAddRmsnorm
>>> 锁定设备 5 并启动算子: 1_RotaryMul
>>> 锁定设备 6 并启动算子: 21_GaussianTopkSparseActivation
>>> 锁定设备 7 并启动算子: 22_HybridAttentionMaskPreparation
Performance Report
========================================================================================
Operator  : /home/y00889327/AscendOpGenAgent/output_0410_l2_ziji/13_InterleaveRope
Task Dir  : /home/y00889327/AscendOpGenAgent/output_0410_l2_ziji/13_InterleaveRope
Device    : npu
Warmup    : 5
Repeat    : 10
Seed      : 0
----------------------------------------------------------------------------------------
Inputs
----------------------------------------------------------------------------------------
inputs[0]: list[3]
inputs[0][0]: Tensor(shape=(1, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[0][1]: Tensor(shape=(1, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[0][2]: Tensor(shape=(1, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[1]: list[3]
inputs[1][0]: Tensor(shape=(1, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[1][1]: Tensor(shape=(1, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[1][2]: Tensor(shape=(1, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[2]: list[3]
inputs[2][0]: Tensor(shape=(1, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[2][1]: Tensor(shape=(1, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[2][2]: Tensor(shape=(1, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[3]: list[3]
inputs[3][0]: Tensor(shape=(1, 1, 256, 64), dtype=torch.float16, device=npu:0)
inputs[3][1]: Tensor(shape=(1, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[3][2]: Tensor(shape=(1, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[4]: list[3]
inputs[4][0]: Tensor(shape=(1, 1, 512, 64), dtype=torch.float16, device=npu:0)
inputs[4][1]: Tensor(shape=(1, 1, 512, 64), dtype=torch.float16, device=npu:0)
inputs[4][2]: Tensor(shape=(1, 1, 512, 64), dtype=torch.float16, device=npu:0)
inputs[5]: list[3]
inputs[5][0]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[5][1]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[5][2]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[6]: list[3]
inputs[6][0]: Tensor(shape=(1, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[6][1]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[6][2]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[7]: list[3]
inputs[7][0]: Tensor(shape=(1, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[7][1]: Tensor(shape=(1, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[7][2]: Tensor(shape=(1, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[8]: list[3]
inputs[8][0]: Tensor(shape=(1, 1, 256, 64), dtype=torch.bfloat16, device=npu:0)
inputs[8][1]: Tensor(shape=(1, 1, 256, 64), dtype=torch.bfloat16, device=npu:0)
inputs[8][2]: Tensor(shape=(1, 1, 256, 64), dtype=torch.bfloat16, device=npu:0)
inputs[9]: list[3]
inputs[9][0]: Tensor(shape=(1, 1, 512, 64), dtype=torch.bfloat16, device=npu:0)
inputs[9][1]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[9][2]: Tensor(shape=(1, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[10]: list[3]
inputs[10][0]: Tensor(shape=(2, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[10][1]: Tensor(shape=(2, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[10][2]: Tensor(shape=(2, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[11]: list[3]
inputs[11][0]: Tensor(shape=(2, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[11][1]: Tensor(shape=(2, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[11][2]: Tensor(shape=(2, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[12]: list[3]
inputs[12][0]: Tensor(shape=(2, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[12][1]: Tensor(shape=(2, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[12][2]: Tensor(shape=(2, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[13]: list[3]
inputs[13][0]: Tensor(shape=(2, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[13][1]: Tensor(shape=(2, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[13][2]: Tensor(shape=(2, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[14]: list[3]
inputs[14][0]: Tensor(shape=(2, 1, 256, 64), dtype=torch.float16, device=npu:0)
inputs[14][1]: Tensor(shape=(2, 1, 256, 64), dtype=torch.float16, device=npu:0)
inputs[14][2]: Tensor(shape=(2, 1, 256, 64), dtype=torch.float16, device=npu:0)
inputs[15]: list[3]
inputs[15][0]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[15][1]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[15][2]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[16]: list[3]
inputs[16][0]: Tensor(shape=(2, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[16][1]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[16][2]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[17]: list[3]
inputs[17][0]: Tensor(shape=(2, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[17][1]: Tensor(shape=(2, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[17][2]: Tensor(shape=(2, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[18]: list[3]
inputs[18][0]: Tensor(shape=(2, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[18][1]: Tensor(shape=(2, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[18][2]: Tensor(shape=(2, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[19]: list[3]
inputs[19][0]: Tensor(shape=(2, 1, 256, 64), dtype=torch.bfloat16, device=npu:0)
inputs[19][1]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[19][2]: Tensor(shape=(2, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[20]: list[3]
inputs[20][0]: Tensor(shape=(4, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[20][1]: Tensor(shape=(4, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[20][2]: Tensor(shape=(4, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[21]: list[3]
inputs[21][0]: Tensor(shape=(4, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[21][1]: Tensor(shape=(4, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[21][2]: Tensor(shape=(4, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[22]: list[3]
inputs[22][0]: Tensor(shape=(4, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[22][1]: Tensor(shape=(4, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[22][2]: Tensor(shape=(4, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[23]: list[3]
inputs[23][0]: Tensor(shape=(4, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[23][1]: Tensor(shape=(4, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[23][2]: Tensor(shape=(4, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[24]: list[3]
inputs[24][0]: Tensor(shape=(4, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[24][1]: Tensor(shape=(4, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[24][2]: Tensor(shape=(4, 1, 128, 64), dtype=torch.float16, device=npu:0)
inputs[25]: list[3]
inputs[25][0]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[25][1]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[25][2]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[26]: list[3]
inputs[26][0]: Tensor(shape=(4, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[26][1]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[26][2]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[27]: list[3]
inputs[27][0]: Tensor(shape=(4, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[27][1]: Tensor(shape=(4, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[27][2]: Tensor(shape=(4, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[28]: list[3]
inputs[28][0]: Tensor(shape=(4, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[28][1]: Tensor(shape=(4, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[28][2]: Tensor(shape=(4, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[29]: list[3]
inputs[29][0]: Tensor(shape=(4, 1, 128, 64), dtype=torch.bfloat16, device=npu:0)
inputs[29][1]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[29][2]: Tensor(shape=(4, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[30]: list[3]
inputs[30][0]: Tensor(shape=(8, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[30][1]: Tensor(shape=(8, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[30][2]: Tensor(shape=(8, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[31]: list[3]
inputs[31][0]: Tensor(shape=(8, 1, 8, 64), dtype=torch.float16, device=npu:0)
inputs[31][1]: Tensor(shape=(8, 1, 8, 64), dtype=torch.float16, device=npu:0)
inputs[31][2]: Tensor(shape=(8, 1, 8, 64), dtype=torch.float16, device=npu:0)
inputs[32]: list[3]
inputs[32][0]: Tensor(shape=(8, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[32][1]: Tensor(shape=(8, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[32][2]: Tensor(shape=(8, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[33]: list[3]
inputs[33][0]: Tensor(shape=(8, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[33][1]: Tensor(shape=(8, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[33][2]: Tensor(shape=(8, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[34]: list[3]
inputs[34][0]: Tensor(shape=(8, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[34][1]: Tensor(shape=(8, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[34][2]: Tensor(shape=(8, 1, 64, 64), dtype=torch.float16, device=npu:0)
inputs[35]: list[3]
inputs[35][0]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[35][1]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[35][2]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[36]: list[3]
inputs[36][0]: Tensor(shape=(8, 1, 8, 64), dtype=torch.bfloat16, device=npu:0)
inputs[36][1]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[36][2]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[37]: list[3]
inputs[37][0]: Tensor(shape=(8, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[37][1]: Tensor(shape=(8, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[37][2]: Tensor(shape=(8, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[38]: list[3]
inputs[38][0]: Tensor(shape=(8, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[38][1]: Tensor(shape=(8, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[38][2]: Tensor(shape=(8, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[39]: list[3]
inputs[39][0]: Tensor(shape=(8, 1, 64, 64), dtype=torch.bfloat16, device=npu:0)
inputs[39][1]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[39][2]: Tensor(shape=(8, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[40]: list[3]
inputs[40][0]: Tensor(shape=(16, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[40][1]: Tensor(shape=(16, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[40][2]: Tensor(shape=(16, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[41]: list[3]
inputs[41][0]: Tensor(shape=(16, 1, 4, 64), dtype=torch.float16, device=npu:0)
inputs[41][1]: Tensor(shape=(16, 1, 4, 64), dtype=torch.float16, device=npu:0)
inputs[41][2]: Tensor(shape=(16, 1, 4, 64), dtype=torch.float16, device=npu:0)
inputs[42]: list[3]
inputs[42][0]: Tensor(shape=(16, 1, 8, 64), dtype=torch.float16, device=npu:0)
inputs[42][1]: Tensor(shape=(16, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[42][2]: Tensor(shape=(16, 1, 1, 64), dtype=torch.float16, device=npu:0)
inputs[43]: list[3]
inputs[43][0]: Tensor(shape=(16, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[43][1]: Tensor(shape=(16, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[43][2]: Tensor(shape=(16, 1, 16, 64), dtype=torch.float16, device=npu:0)
inputs[44]: list[3]
inputs[44][0]: Tensor(shape=(16, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[44][1]: Tensor(shape=(16, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[44][2]: Tensor(shape=(16, 1, 32, 64), dtype=torch.float16, device=npu:0)
inputs[45]: list[3]
inputs[45][0]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[45][1]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[45][2]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[46]: list[3]
inputs[46][0]: Tensor(shape=(16, 1, 4, 64), dtype=torch.bfloat16, device=npu:0)
inputs[46][1]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[46][2]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[47]: list[3]
inputs[47][0]: Tensor(shape=(16, 1, 8, 64), dtype=torch.bfloat16, device=npu:0)
inputs[47][1]: Tensor(shape=(16, 1, 8, 64), dtype=torch.bfloat16, device=npu:0)
inputs[47][2]: Tensor(shape=(16, 1, 8, 64), dtype=torch.bfloat16, device=npu:0)
inputs[48]: list[3]
inputs[48][0]: Tensor(shape=(16, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[48][1]: Tensor(shape=(16, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[48][2]: Tensor(shape=(16, 1, 16, 64), dtype=torch.bfloat16, device=npu:0)
inputs[49]: list[3]
inputs[49][0]: Tensor(shape=(16, 1, 32, 64), dtype=torch.bfloat16, device=npu:0)
inputs[49][1]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
inputs[49][2]: Tensor(shape=(16, 1, 1, 64), dtype=torch.bfloat16, device=npu:0)
----------------------------------------------------------------------------------------
Impl         Status       Mean(ms)       Median          Min          Max          Std
----------------------------------------------------------------------------------------
reference    OK              0.080        0.044        0.029        5.230        0.345
tilelang     OK              0.269        0.253        0.225        1.091        0.066
ascendc      OK              0.187        0.174        0.153        0.912        0.066
----------------------------------------------------------------------------------------
reference -> /home/y00889327/AscendOpGenAgent/output_0410_l2_ziji/13_InterleaveRope/model.py
case[0] mean=0.062 ms, median=0.055 ms, samples(ms): [0.109, 0.055, 0.062, 0.055, 0.052, 0.057, 0.084, 0.055, 0.046, 0.045]
case[1] mean=0.044 ms, median=0.045 ms, samples(ms): [0.046, 0.044, 0.047, 0.044, 0.037, 0.045, 0.039, 0.045, 0.050, 0.043]
case[2] mean=0.049 ms, median=0.048 ms, samples(ms): [0.050, 0.053, 0.044, 0.068, 0.048, 0.049, 0.039, 0.047, 0.045, 0.047]
case[3] mean=0.340 ms, median=0.049 ms, samples(ms): [0.044, 2.832, 0.162, 0.056, 0.043, 0.073, 0.050, 0.044, 0.045, 0.047]
case[4] mean=0.044 ms, median=0.044 ms, samples(ms): [0.042, 0.048, 0.045, 0.044, 0.041, 0.043, 0.045, 0.041, 0.046, 0.040]
case[5] mean=0.039 ms, median=0.039 ms, samples(ms): [0.042, 0.039, 0.038, 0.040, 0.039, 0.037, 0.037, 0.036, 0.039, 0.043]
case[6] mean=0.042 ms, median=0.041 ms, samples(ms): [0.041, 0.041, 0.045, 0.042, 0.046, 0.046, 0.040, 0.044, 0.040, 0.038]
case[7] mean=0.043 ms, median=0.042 ms, samples(ms): [0.042, 0.037, 0.043, 0.061, 0.040, 0.045, 0.040, 0.043, 0.042, 0.039]
case[8] mean=0.043 ms, median=0.043 ms, samples(ms): [0.044, 0.040, 0.046, 0.047, 0.040, 0.044, 0.040, 0.047, 0.040, 0.042]
case[9] mean=0.044 ms, median=0.043 ms, samples(ms): [0.048, 0.045, 0.044, 0.045, 0.042, 0.042, 0.043, 0.049, 0.041, 0.041]
case[10] mean=0.039 ms, median=0.038 ms, samples(ms): [0.044, 0.038, 0.037, 0.035, 0.037, 0.035, 0.043, 0.042, 0.039, 0.040]
case[11] mean=0.042 ms, median=0.042 ms, samples(ms): [0.044, 0.037, 0.047, 0.040, 0.039, 0.045, 0.043, 0.043, 0.042, 0.037]
case[12] mean=0.045 ms, median=0.043 ms, samples(ms): [0.045, 0.042, 0.040, 0.058, 0.041, 0.045, 0.043, 0.048, 0.042, 0.043]
case[13] mean=0.042 ms, median=0.042 ms, samples(ms): [0.046, 0.045, 0.039, 0.046, 0.041, 0.043, 0.048, 0.039, 0.038, 0.039]
case[14] mean=0.043 ms, median=0.042 ms, samples(ms): [0.040, 0.044, 0.042, 0.040, 0.050, 0.040, 0.041, 0.044, 0.041, 0.045]
case[15] mean=0.041 ms, median=0.041 ms, samples(ms): [0.045, 0.041, 0.041, 0.044, 0.042, 0.042, 0.039, 0.037, 0.038, 0.043]
case[16] mean=0.045 ms, median=0.044 ms, samples(ms): [0.039, 0.041, 0.040, 0.046, 0.043, 0.049, 0.048, 0.051, 0.045, 0.041]
case[17] mean=0.044 ms, median=0.042 ms, samples(ms): [0.056, 0.039, 0.048, 0.042, 0.042, 0.039, 0.040, 0.042, 0.044, 0.043]
case[18] mean=0.044 ms, median=0.044 ms, samples(ms): [0.047, 0.050, 0.041, 0.044, 0.045, 0.040, 0.042, 0.046, 0.046, 0.042]
case[19] mean=0.047 ms, median=0.045 ms, samples(ms): [0.056, 0.045, 0.042, 0.044, 0.040, 0.047, 0.056, 0.045, 0.052, 0.045]
case[20] mean=0.045 ms, median=0.042 ms, samples(ms): [0.046, 0.038, 0.056, 0.044, 0.030, 0.039, 0.043, 0.040, 0.072, 0.039]
case[21] mean=0.041 ms, median=0.039 ms, samples(ms): [0.047, 0.037, 0.044, 0.036, 0.037, 0.036, 0.037, 0.043, 0.051, 0.041]
case[22] mean=0.050 ms, median=0.046 ms, samples(ms): [0.056, 0.049, 0.080, 0.041, 0.046, 0.047, 0.042, 0.045, 0.055, 0.042]
case[23] mean=0.043 ms, median=0.041 ms, samples(ms): [0.041, 0.042, 0.039, 0.039, 0.032, 0.040, 0.045, 0.049, 0.050, 0.049]
case[24] mean=0.055 ms, median=0.049 ms, samples(ms): [0.049, 0.057, 0.047, 0.094, 0.052, 0.048, 0.044, 0.046, 0.069, 0.044]
case[25] mean=0.050 ms, median=0.047 ms, samples(ms): [0.072, 0.048, 0.047, 0.044, 0.043, 0.056, 0.047, 0.044, 0.059, 0.043]
case[26] mean=0.050 ms, median=0.044 ms, samples(ms): [0.058, 0.044, 0.059, 0.043, 0.042, 0.051, 0.041, 0.041, 0.076, 0.043]
case[27] mean=0.049 ms, median=0.049 ms, samples(ms): [0.056, 0.049, 0.050, 0.047, 0.045, 0.049, 0.043, 0.054, 0.059, 0.041]
case[28] mean=0.049 ms, median=0.046 ms, samples(ms): [0.070, 0.055, 0.047, 0.043, 0.042, 0.054, 0.044, 0.044, 0.049, 0.044]
case[29] mean=0.052 ms, median=0.051 ms, samples(ms): [0.051, 0.055, 0.046, 0.050, 0.052, 0.053, 0.050, 0.052, 0.066, 0.046]
case[30] mean=0.074 ms, median=0.057 ms, samples(ms): [0.049, 0.051, 0.045, 0.120, 0.176, 0.062, 0.052, 0.061, 0.073, 0.050]
case[31] mean=0.051 ms, median=0.050 ms, samples(ms): [0.063, 0.062, 0.057, 0.052, 0.044, 0.047, 0.043, 0.045, 0.049, 0.050]
case[32] mean=0.051 ms, median=0.048 ms, samples(ms): [0.058, 0.068, 0.043, 0.045, 0.057, 0.047, 0.052, 0.043, 0.050, 0.046]
case[33] mean=0.052 ms, median=0.046 ms, samples(ms): [0.044, 0.055, 0.046, 0.044, 0.055, 0.044, 0.058, 0.044, 0.081, 0.047]
case[34] mean=0.053 ms, median=0.051 ms, samples(ms): [0.055, 0.059, 0.056, 0.068, 0.047, 0.046, 0.047, 0.052, 0.050, 0.047]
case[35] mean=0.052 ms, median=0.048 ms, samples(ms): [0.047, 0.054, 0.043, 0.079, 0.043, 0.043, 0.071, 0.042, 0.049, 0.050]
case[36] mean=1.239 ms, median=0.098 ms, samples(ms): [0.056, 5.230, 0.140, 0.046, 0.040, 2.910, 0.270, 0.044, 0.034, 3.622]
case[37] mean=0.045 ms, median=0.043 ms, samples(ms): [0.045, 0.063, 0.043, 0.044, 0.040, 0.050, 0.040, 0.041, 0.040, 0.047]
case[38] mean=0.046 ms, median=0.041 ms, samples(ms): [0.038, 0.049, 0.068, 0.040, 0.040, 0.040, 0.042, 0.038, 0.051, 0.050]
case[39] mean=0.050 ms, median=0.050 ms, samples(ms): [0.048, 0.053, 0.051, 0.050, 0.050, 0.044, 0.052, 0.045, 0.053, 0.049]
case[40] mean=0.038 ms, median=0.037 ms, samples(ms): [0.038, 0.040, 0.042, 0.037, 0.036, 0.036, 0.037, 0.033, 0.032, 0.045]
case[41] mean=0.038 ms, median=0.037 ms, samples(ms): [0.040, 0.040, 0.036, 0.035, 0.034, 0.035, 0.037, 0.037, 0.036, 0.046]
case[42] mean=0.040 ms, median=0.038 ms, samples(ms): [0.041, 0.046, 0.038, 0.038, 0.035, 0.038, 0.038, 0.039, 0.037, 0.047]
case[43] mean=0.252 ms, median=0.040 ms, samples(ms): [0.040, 0.051, 0.040, 0.040, 0.039, 0.038, 0.039, 0.048, 2.124, 0.058]
case[44] mean=0.044 ms, median=0.041 ms, samples(ms): [0.044, 0.050, 0.041, 0.042, 0.040, 0.041, 0.040, 0.041, 0.040, 0.055]
case[45] mean=0.045 ms, median=0.044 ms, samples(ms): [0.043, 0.041, 0.054, 0.041, 0.042, 0.045, 0.046, 0.040, 0.046, 0.050]
case[46] mean=0.043 ms, median=0.041 ms, samples(ms): [0.044, 0.043, 0.050, 0.041, 0.040, 0.041, 0.041, 0.040, 0.040, 0.047]
case[47] mean=0.047 ms, median=0.041 ms, samples(ms): [0.043, 0.041, 0.049, 0.041, 0.040, 0.040, 0.041, 0.039, 0.047, 0.086]
case[48] mean=0.035 ms, median=0.035 ms, samples(ms): [0.038, 0.037, 0.040, 0.029, 0.037, 0.034, 0.032, 0.032, 0.034, 0.037]
case[49] mean=0.048 ms, median=0.047 ms, samples(ms): [0.046, 0.046, 0.053, 0.044, 0.048, 0.047, 0.051, 0.052, 0.046, 0.050]
----------------------------------------------------------------------------------------
tilelang -> /home/y00889327/AscendOpGenAgent/output_0410_l2_ziji/13_InterleaveRope/model_new_tilelang.py
case[0] mean=0.499 ms, median=0.360 ms, samples(ms): [0.358, 0.331, 0.379, 1.091, 0.404, 0.362, 0.347, 0.334, 0.350, 1.032]
case[1] mean=0.382 ms, median=0.396 ms, samples(ms): [0.351, 0.344, 0.336, 0.349, 0.418, 0.397, 0.428, 0.399, 0.395, 0.401]
case[2] mean=0.372 ms, median=0.378 ms, samples(ms): [0.321, 0.323, 0.319, 0.362, 0.365, 0.402, 0.392, 0.393, 0.446, 0.393]
case[3] mean=0.433 ms, median=0.434 ms, samples(ms): [0.457, 0.505, 0.389, 0.370, 0.443, 0.425, 0.423, 0.448, 0.420, 0.449]
case[4] mean=0.259 ms, median=0.252 ms, samples(ms): [0.253, 0.246, 0.241, 0.256, 0.257, 0.253, 0.248, 0.252, 0.248, 0.340]
case[5] mean=0.232 ms, median=0.231 ms, samples(ms): [0.234, 0.235, 0.225, 0.233, 0.230, 0.227, 0.241, 0.229, 0.236, 0.226]
case[6] mean=0.278 ms, median=0.271 ms, samples(ms): [0.278, 0.266, 0.272, 0.260, 0.259, 0.334, 0.277, 0.270, 0.268, 0.297]
case[7] mean=0.277 ms, median=0.264 ms, samples(ms): [0.267, 0.280, 0.283, 0.351, 0.321, 0.257, 0.251, 0.262, 0.252, 0.241]
case[8] mean=0.251 ms, median=0.253 ms, samples(ms): [0.254, 0.241, 0.262, 0.246, 0.257, 0.253, 0.243, 0.262, 0.252, 0.242]
case[9] mean=0.265 ms, median=0.265 ms, samples(ms): [0.266, 0.262, 0.266, 0.256, 0.278, 0.270, 0.257, 0.274, 0.264, 0.261]
case[10] mean=0.252 ms, median=0.251 ms, samples(ms): [0.259, 0.257, 0.254, 0.248, 0.254, 0.249, 0.247, 0.252, 0.247, 0.251]
case[11] mean=0.249 ms, median=0.249 ms, samples(ms): [0.249, 0.251, 0.252, 0.254, 0.245, 0.249, 0.245, 0.242, 0.255, 0.244]
case[12] mean=0.299 ms, median=0.264 ms, samples(ms): [0.270, 0.261, 0.264, 0.263, 0.257, 0.608, 0.271, 0.256, 0.265, 0.272]
case[13] mean=0.252 ms, median=0.249 ms, samples(ms): [0.248, 0.252, 0.272, 0.246, 0.247, 0.248, 0.252, 0.258, 0.251, 0.245]
case[14] mean=0.251 ms, median=0.250 ms, samples(ms): [0.252, 0.250, 0.252, 0.247, 0.251, 0.247, 0.246, 0.262, 0.249, 0.256]
case[15] mean=0.251 ms, median=0.250 ms, samples(ms): [0.256, 0.248, 0.247, 0.251, 0.258, 0.245, 0.249, 0.255, 0.248, 0.251]
case[16] mean=0.267 ms, median=0.267 ms, samples(ms): [0.261, 0.272, 0.258, 0.278, 0.267, 0.256, 0.270, 0.280, 0.257, 0.268]
case[17] mean=0.264 ms, median=0.264 ms, samples(ms): [0.262, 0.248, 0.267, 0.254, 0.248, 0.274, 0.274, 0.296, 0.268, 0.249]
case[18] mean=0.251 ms, median=0.248 ms, samples(ms): [0.255, 0.250, 0.266, 0.256, 0.257, 0.246, 0.247, 0.242, 0.246, 0.246]
case[19] mean=0.264 ms, median=0.262 ms, samples(ms): [0.267, 0.279, 0.261, 0.271, 0.268, 0.255, 0.262, 0.263, 0.257, 0.260]
case[20] mean=0.252 ms, median=0.251 ms, samples(ms): [0.254, 0.258, 0.248, 0.256, 0.271, 0.241, 0.259, 0.241, 0.247, 0.243]
case[21] mean=0.250 ms, median=0.250 ms, samples(ms): [0.251, 0.249, 0.255, 0.254, 0.248, 0.252, 0.254, 0.246, 0.244, 0.245]
case[22] mean=0.263 ms, median=0.262 ms, samples(ms): [0.262, 0.261, 0.269, 0.263, 0.261, 0.264, 0.262, 0.261, 0.265, 0.258]
case[23] mean=0.254 ms, median=0.252 ms, samples(ms): [0.268, 0.259, 0.253, 0.249, 0.258, 0.252, 0.247, 0.251, 0.259, 0.248]
case[24] mean=0.248 ms, median=0.248 ms, samples(ms): [0.248, 0.254, 0.249, 0.253, 0.246, 0.242, 0.255, 0.243, 0.251, 0.243]
case[25] mean=0.245 ms, median=0.244 ms, samples(ms): [0.251, 0.244, 0.248, 0.244, 0.249, 0.244, 0.245, 0.250, 0.241, 0.238]
case[26] mean=0.258 ms, median=0.255 ms, samples(ms): [0.270, 0.262, 0.254, 0.268, 0.255, 0.256, 0.253, 0.253, 0.255, 0.255]
case[27] mean=0.241 ms, median=0.241 ms, samples(ms): [0.243, 0.244, 0.240, 0.239, 0.242, 0.244, 0.238, 0.243, 0.239, 0.237]
case[28] mean=0.251 ms, median=0.244 ms, samples(ms): [0.244, 0.243, 0.242, 0.243, 0.247, 0.243, 0.308, 0.253, 0.246, 0.244]
case[29] mean=0.256 ms, median=0.254 ms, samples(ms): [0.263, 0.257, 0.259, 0.261, 0.254, 0.252, 0.249, 0.254, 0.254, 0.251]
case[30] mean=0.243 ms, median=0.243 ms, samples(ms): [0.250, 0.245, 0.242, 0.247, 0.245, 0.243, 0.240, 0.240, 0.239, 0.241]
case[31] mean=0.244 ms, median=0.243 ms, samples(ms): [0.255, 0.244, 0.239, 0.243, 0.247, 0.243, 0.240, 0.244, 0.242, 0.244]
case[32] mean=0.259 ms, median=0.259 ms, samples(ms): [0.258, 0.262, 0.265, 0.253, 0.263, 0.256, 0.260, 0.261, 0.256, 0.252]
case[33] mean=0.241 ms, median=0.243 ms, samples(ms): [0.247, 0.243, 0.243, 0.235, 0.252, 0.243, 0.238, 0.234, 0.235, 0.243]
case[34] mean=0.241 ms, median=0.240 ms, samples(ms): [0.242, 0.240, 0.246, 0.243, 0.244, 0.239, 0.238, 0.241, 0.240, 0.236]
case[35] mean=0.246 ms, median=0.245 ms, samples(ms): [0.257, 0.250, 0.243, 0.255, 0.241, 0.239, 0.248, 0.242, 0.246, 0.241]
case[36] mean=0.258 ms, median=0.255 ms, samples(ms): [0.262, 0.254, 0.254, 0.255, 0.265, 0.265, 0.253, 0.253, 0.254, 0.266]
case[37] mean=0.253 ms, median=0.246 ms, samples(ms): [0.302, 0.265, 0.246, 0.246, 0.248, 0.246, 0.244, 0.246, 0.247, 0.244]
case[38] mean=0.245 ms, median=0.244 ms, samples(ms): [0.250, 0.240, 0.243, 0.249, 0.247, 0.246, 0.242, 0.247, 0.243, 0.242]
case[39] mean=0.259 ms, median=0.260 ms, samples(ms): [0.266, 0.252, 0.262, 0.255, 0.256, 0.256, 0.261, 0.264, 0.259, 0.262]
case[40] mean=0.248 ms, median=0.247 ms, samples(ms): [0.248, 0.246, 0.254, 0.248, 0.239, 0.252, 0.247, 0.243, 0.254, 0.247]
case[41] mean=0.246 ms, median=0.244 ms, samples(ms): [0.251, 0.250, 0.259, 0.248, 0.238, 0.240, 0.243, 0.243, 0.239, 0.245]
case[42] mean=0.259 ms, median=0.257 ms, samples(ms): [0.258, 0.257, 0.270, 0.255, 0.256, 0.255, 0.264, 0.256, 0.252, 0.272]
case[43] mean=0.248 ms, median=0.246 ms, samples(ms): [0.246, 0.252, 0.246, 0.245, 0.248, 0.242, 0.247, 0.243, 0.245, 0.264]
case[44] mean=0.251 ms, median=0.247 ms, samples(ms): [0.245, 0.259, 0.246, 0.245, 0.248, 0.250, 0.243, 0.242, 0.265, 0.265]
case[45] mean=0.246 ms, median=0.246 ms, samples(ms): [0.245, 0.247, 0.248, 0.247, 0.248, 0.242, 0.242, 0.244, 0.245, 0.254]
case[46] mean=0.257 ms, median=0.255 ms, samples(ms): [0.258, 0.255, 0.256, 0.254, 0.254, 0.261, 0.253, 0.259, 0.271, 0.253]
case[47] mean=0.315 ms, median=0.314 ms, samples(ms): [0.321, 0.313, 0.313, 0.313, 0.320, 0.311, 0.313, 0.322, 0.314, 0.314]
case[48] mean=0.261 ms, median=0.249 ms, samples(ms): [0.319, 0.298, 0.250, 0.257, 0.248, 0.246, 0.245, 0.247, 0.243, 0.257]
case[49] mean=0.262 ms, median=0.260 ms, samples(ms): [0.260, 0.260, 0.269, 0.262, 0.259, 0.257, 0.257, 0.259, 0.274, 0.261]
----------------------------------------------------------------------------------------
ascendc -> /home/y00889327/AscendOpGenAgent/output_0410_l2_ziji/13_InterleaveRope/model_new_ascendc.py
case[0] mean=0.166 ms, median=0.162 ms, samples(ms): [0.174, 0.180, 0.176, 0.166, 0.163, 0.159, 0.162, 0.158, 0.159, 0.162]
case[1] mean=0.170 ms, median=0.169 ms, samples(ms): [0.177, 0.169, 0.176, 0.174, 0.173, 0.165, 0.169, 0.166, 0.167, 0.167]
case[2] mean=0.176 ms, median=0.174 ms, samples(ms): [0.175, 0.173, 0.185, 0.172, 0.190, 0.177, 0.172, 0.175, 0.167, 0.171]
case[3] mean=0.183 ms, median=0.182 ms, samples(ms): [0.184, 0.190, 0.182, 0.183, 0.179, 0.178, 0.177, 0.188, 0.182, 0.188]
case[4] mean=0.173 ms, median=0.170 ms, samples(ms): [0.178, 0.169, 0.170, 0.174, 0.171, 0.173, 0.190, 0.170, 0.168, 0.169]
case[5] mean=0.164 ms, median=0.159 ms, samples(ms): [0.170, 0.174, 0.188, 0.167, 0.159, 0.156, 0.158, 0.157, 0.153, 0.157]
case[6] mean=0.184 ms, median=0.182 ms, samples(ms): [0.199, 0.186, 0.183, 0.187, 0.181, 0.183, 0.182, 0.182, 0.181, 0.178]
case[7] mean=0.173 ms, median=0.170 ms, samples(ms): [0.190, 0.175, 0.169, 0.168, 0.168, 0.183, 0.171, 0.170, 0.172, 0.165]
case[8] mean=0.173 ms, median=0.171 ms, samples(ms): [0.179, 0.173, 0.177, 0.171, 0.170, 0.171, 0.170, 0.178, 0.171, 0.168]
case[9] mean=0.186 ms, median=0.186 ms, samples(ms): [0.190, 0.187, 0.187, 0.185, 0.182, 0.183, 0.185, 0.188, 0.186, 0.185]
case[10] mean=0.186 ms, median=0.176 ms, samples(ms): [0.185, 0.176, 0.176, 0.171, 0.172, 0.171, 0.172, 0.198, 0.220, 0.218]
case[11] mean=0.215 ms, median=0.212 ms, samples(ms): [0.217, 0.239, 0.216, 0.211, 0.215, 0.212, 0.211, 0.209, 0.211, 0.211]
case[12] mean=0.199 ms, median=0.192 ms, samples(ms): [0.226, 0.227, 0.220, 0.190, 0.198, 0.187, 0.193, 0.188, 0.181, 0.179]
case[13] mean=0.170 ms, median=0.170 ms, samples(ms): [0.170, 0.172, 0.169, 0.169, 0.166, 0.17

@Just-it

Just-it commented Apr 13, 2026

Copy link
Copy Markdown
Owner

1、验证移到skill里面
2、缺少对torch_npu的约束
3、这两个脚本是可以归一的
参照一下这个PR:#81

@chopper0126 chopper0126 changed the title 添加回退校验+迭代循环 添加回退校验+迭代循环+[修改input_case 为从json读取] Apr 13, 2026
@Just-it

Just-it commented Apr 16, 2026 via email

Copy link
Copy Markdown
Owner

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants