Where every model stands
On tool calling we are 10.5 points behind LFM2.5's best 4-bit build on single-turn BFCL and score 0 on multi-turn. Our misses are mostly wrong values in the right slots (wrong-slot rate 10.8-17.8% vs LFM 5.3%). On chat, v2 roughly doubled IFEval over v1 but stays at about half of LFM2.5, and half of its passes loop. Decode speed on the M4 GPU now beats LFM2.5; prefill on long prompts does not.
111 models from the first training on the H200 node (1 October, 13:10 IST) to 6 October ~09:00 IST. Times are IST. BFCL numbers are marked S (single-turn) or O (official, multi-turn counted as zero).
Best model per area, against LFM2.5
Tool calling, BFCL single-turn
Tool calling, rows right
Tool calling, multi-turn
Tool calling, as a pipeline
Chat / instruction following
Chat, refusals on benign prompts
Picker (general "pick from options")
Retrieval (film database)
Tonttu app (intent + args)
Specialists, held-out gates
Title (editor read, 100 rows)
On-device speed (the M4 Mac mini M4, 64/512/2000-token prompts)
Size on device
Running now and next
| what | owner | where | state | ETA (IST) |
|---|---|---|---|---|
| G-453 general picker, shard 2 scoring (suite, BFCL behind 3 choosers, 18-family held-out) | node GPUs 4-7 (4 eval processes) | shard 2 trained 06:42-08:23, | scored ~08:48 | |
| G-453 shard 3 training (2.2M rows, 1.003B tokens, last part of one cosine) | node GPUs 4-7 | gated on shard 2 scoring | trained ~11:08, GPUs released ~11:35 | |
| G-455 chat+tool SFT v3 (over-refusal fix, 1,475.9M tokens, 2,815 steps) | node GPUs 4-7 | armed on "G-453 COMPLETE run=picker-3b … gpus=4-7 released" | after ~11:35; SCORED was ~11:15, now later | |
| G-456 general retrieval reader data (Luna writes 20,900 answers, cap) | node CPU (build_reader.py running) | data build | not trained yet | |
| GPUs 0-3 | empty (0 MiB), as the user's rule requires |
Before you read the numbers
added). Every count carries its denominator.
Window. Section 0 (added for G-460) starts at our first training job on the H200 node, 2026-10-01 13:10 IST, and covers 10-01 to 10-03. Sections 1-7 cover models whose training ran, ended or was first scored between 2026-10-04 00:00 and 2026-10-06 ~09:00 IST. The h1280 200B base started on 10-03 and ended in the window, so it is included. Section 3.10 lists two models trained just before the window that were accepted inside it.
Two BFCL scales. Read this first. BFCL v3 numbers in the project come on two scales, and the same model gets both:
- S (single-turn, "without multi-turn"): the mean over the 13 single-turn categories, 3,637 prompts (3,641 rows).
The user asked on 10-05 for BFCL to be reported this way. All numbers after 10-05 ~04:40 IST are on S.
- O (official formula): multi-turn counted as 0, so O ≈ S × 2/3. Most numbers before 10-05 04:40 IST are on O.
Examples of the same model on both: v2 per-part RL s200 42.96 S = 28.64 O; v2 final 38.70 S = 25.80 O; old wide pack 0.5B 25.08 S = 16.72 O; Jev picker 56.35 S = 37.57 O; LFM2.5 MLX 4-bit 42.82 S = 28.54 O.
- Only LFM2.5 QAD-Q4_0 has a real multi-turn score: 4/800, official overall 35.78. Our best pack scores 0/800.
- means no schema-constrained decoding. The user ruled schema rules out as an evaluation (10-05 night).
Hardware names. "node" = the shared 8x H200 box ( from 10-05 evening only GPUs 4-7 are ours (user rule), with a one-off loan of GPUs 0-3 from 02:46 to 06:00 IST on 10-06. and are single rented Lambda boxes. "M1 ANE" = the M1 MacBook's Neural Engine (the M1 MacBook); "the M4 Mac mini" = the M4 Mac mini.