Where every model stands

On tool calling we are 10.5 points behind LFM2.5's best 4-bit build on single-turn BFCL and score 0 on multi-turn. Our misses are mostly wrong values in the right slots (wrong-slot rate 10.8-17.8% vs LFM 5.3%). On chat, v2 roughly doubled IFEval over v1 but stays at about half of LFM2.5, and half of its passes loop. Decode speed on the M4 GPU now beats LFM2.5; prefill on long prompts does not.

111 models from the first training on the H200 node (1 October, 13:10 IST) to 6 October ~09:00 IST. Times are IST. BFCL numbers are marked S (single-turn) or O (official, multi-turn counted as zero).

Best model per area, against LFM2.5

Tool calling, BFCL single-turn

v2 per-part RL step 200
Ours42.96 S (official 28.64; one M1 ANE run, no plain GPU run exists)
LFM2.5-230M or referenceQAD-Q4_0 53.42 S (official 35.78); MLX 4-bit 42.82 S

Tool calling, rows right

v2 per-part RL step 200
Ours1,617 / 3,641
LFM2.5-230M or referenceQAD 1,994 / 3,641 (ours only 348, LFM only 725, McNemar p=4e-31)

Tool calling, multi-turn

v2 per-part RL step 200
Ours0 / 800
LFM2.5-230M or referenceQAD 4 / 800

Tool calling, as a pipeline

picker G-453 shard 1 filling args behind the pack's tool choice
Ours43.02 S (behind gold tools 53.64, behind Jev's choices 43.96)
LFM2.5-230M or referenceNot run on LFM2.5. Reference: the pack alone 42.96

Chat / instruction following

chat+tool SFT v2
OursIFEval strict 180/541 (89 of the 180 not looping); IFBench 34/300; Multi-IF turns 1/2/3 strict 271/78/23 (overall 0.335/0.220/0.172)
LFM2.5-230M or referenceIFEval 329/541 (308 not looping by the owner's count; 313/328 by the Manager's); IFBench 109/300; Multi-IF 440/203/66 (0.507/0.369/0.274)

Chat, refusals on benign prompts

chat+tool SFT v2
Ours11/100 real100, 11/541 IFEval (v1: 3 and 3)
LFM2.5-230M or referenceIFEval refusals 4/541

Picker (general "pick from options")

G-453 shard 1
Oursheld-out 18 families: 1,974 / 3,600; schema-given suite 4,374 / 5,030 fields
LFM2.5-230M or referenceNot run on LFM2.5. Reference: Jev 1.13: 3,050 / 3,600 and 3,951 / 5,030; chance 507 / 3,600

Retrieval (film database)

retrieval pack v2
Oursfetch@1 1,102 / 1,153 (BM25@1 915); end to end 1,005 / 1,153
LFM2.5-230M or referenceNot run on LFM2.5. Reference: redefined at 08:11 IST as a general reader (G-456), not trained yet

Tonttu app (intent + args)

unified Tonttu pack, seed 20260921
Oursintent 430/440, intent+args 358/440; new 92: 92/92 and 88/92
LFM2.5-230M or referenceNot run on LFM2.5. Reference: Tonttu's own 26 Core ML models 398/310

Specialists, held-out gates

uhm filler tagger
Ourspasses 4/4 gate lines
LFM2.5-230M or referenceNot run on LFM2.5. Reference: address field 90.3% passes; redact TAB 862/987 passes with the merge rule; spam, moderator, receipts, title, gist, docqa fail; thread read owed (parked)

Title (editor read, 100 rows)

title v3 data pack
Oursaccurate and plain 8/100, invented facts 88/100
LFM2.5-230M or reference17/100 and 61/100

On-device speed (the M4 Mac mini M4, 64/512/2000-token prompts)

ternary Metal engine (runtime/ternary-metal,)
Oursdecode 697/657/581 tok/s; prefill 4,407/5,716/5,366
LFM2.5-230M or referenceMLX 4-bit decode 563/525/390; prefill 4,762/7,784/9,130

Size on device

wide pack, palettized 2-bit
Ours113 MB
LFM2.5-230M or reference139 MB

Running now and next

whatownerwherestateETA (IST)
G-453 general picker, shard 2 scoring (suite, BFCL behind 3 choosers, 18-family held-out)node GPUs 4-7 (4 eval processes)shard 2 trained 06:42-08:23,scored ~08:48
G-453 shard 3 training (2.2M rows, 1.003B tokens, last part of one cosine)node GPUs 4-7gated on shard 2 scoringtrained ~11:08, GPUs released ~11:35
G-455 chat+tool SFT v3 (over-refusal fix, 1,475.9M tokens, 2,815 steps)node GPUs 4-7armed on "G-453 COMPLETE run=picker-3b … gpus=4-7 released"after ~11:35; SCORED was ~11:15, now later
G-456 general retrieval reader data (Luna writes 20,900 answers, cap)node CPU (build_reader.py running)data buildnot trained yet
GPUs 0-3empty (0 MiB), as the user's rule requires

Before you read the numbers

added). Every count carries its denominator.

Window. Section 0 (added for G-460) starts at our first training job on the H200 node, 2026-10-01 13:10 IST, and covers 10-01 to 10-03. Sections 1-7 cover models whose training ran, ended or was first scored between 2026-10-04 00:00 and 2026-10-06 ~09:00 IST. The h1280 200B base started on 10-03 and ended in the window, so it is included. Section 3.10 lists two models trained just before the window that were accepted inside it.

Two BFCL scales. Read this first. BFCL v3 numbers in the project come on two scales, and the same model gets both:

The user asked on 10-05 for BFCL to be reported this way. All numbers after 10-05 ~04:40 IST are on S.

Examples of the same model on both: v2 per-part RL s200 42.96 S = 28.64 O; v2 final 38.70 S = 25.80 O; old wide pack 0.5B 25.08 S = 16.72 O; Jev picker 56.35 S = 37.57 O; LFM2.5 MLX 4-bit 42.82 S = 28.54 O.

Hardware names. "node" = the shared 8x H200 box ( from 10-05 evening only GPUs 4-7 are ours (user rule), with a one-off loan of GPUs 0-3 from 02:46 to 06:00 IST on 10-06. and are single rented Lambda boxes. "M1 ANE" = the M1 MacBook's Neural Engine (the M1 MacBook); "the M4 Mac mini" = the M4 Mac mini.