Skip to content

Instantly share code, notes, and snippets.

@hizkifw
Created July 13, 2026 06:47
Show Gist options
  • Select an option

  • Save hizkifw/c7f19222c71c4a33545611bf43108f03 to your computer and use it in GitHub Desktop.

Select an option

Save hizkifw/c7f19222c71c4a33545611bf43108f03 to your computer and use it in GitHub Desktop.

ThinkingCap vs. Qwen3.6-27B Abliterated: Reasoning Token Benchmark

This is a small paired benchmark comparing the reasoning-token usage of:

Both models are dense Qwen3.6 27B derivatives. The Q4_K label is an alias of the Q4_K_M mixed quantization recipe in current llama.cpp, and both files were approximately 16.81 GB.

Test setup

  • GPU: NVIDIA GeForce RTX 3090, 24 GB
  • Runtime: LM Studio / llama.cpp
  • Context length: 32,768 tokens
  • GPU offload: 100%
  • Parallel slots: 1
  • Native MTP speculative decoding: enabled
  • Maximum speculative tokens: 3
  • Temperature: 0
  • Maximum completion tokens: 4,096
  • Runs per prompt: 1

Reasoning-token counts were taken directly from LM Studio's OpenAI-compatible API at usage.completion_tokens_details.reasoning_tokens. Generation time covers the completion request, not model loading.

Summary

Metric ThinkingCap Baseline Difference
Total reasoning tokens 7,028 12,581 44.1% fewer
Mean reasoning tokens 702.8 1,258.1 44.1% fewer
Median reasoning tokens 547.5 1,008 45.7% fewer
Total generation time 132.3 s 230.6 s 42.6% faster
Correct final answers 10/10 9/10 +1
MTP draft acceptance 94.3% 96.5% -2.2 pp

ThinkingCap used fewer reasoning tokens on every prompt. It returned the same correct final answer as the baseline on nine prompts. On the remaining task-ordering prompt, ThinkingCap answered correctly while the baseline exhausted its 4,096-token completion limit without producing a final answer. Because that baseline result was capped, the overall 44.1% token reduction is a conservative figure.

Per-prompt results

Prompt ThinkingCap reasoning Baseline reasoning Reduction ThinkingCap final Baseline final
Discount and tax 541 693 21.9% $220.32 $220.32
Mislabeled boxes 631 978 35.5% Mixed Mixed
Meeting time 667 1,204 44.6% 11:48 AM 11:48 AM
Coin probability 439 743 40.9% 5/16 5/16
Python mutation trace 756 1,629 53.6% 15 15
Age algebra 554 1,038 46.6% 30 30
Task ordering 2,196 3,992 at least 45.0% 5 No final answer; capped
Rectangle area 428 692 38.2% 230 230
Calendar reasoning 438 1,117 60.8% Thursday Thursday
Chickens and rabbits 378 495 23.6% 15 15

Original prompts and detailed results

1. Discount and tax

A store discounts a $240 item by 15%, then applies 8% sales tax to the discounted price. Think carefully and give only the final price in dollars with two decimal places.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 541 551 9.813 s $220.32 Yes
Baseline 693 703 12.388 s $220.32 Yes

2. Mislabeled boxes

Three boxes are labeled Apples, Oranges, and Mixed. Every label is wrong. You may draw one fruit from one box to determine all contents. Think carefully and give only the label of the box you should draw from first.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 631 635 12.956 s Mixed Yes
Baseline 978 982 18.796 s Mixed Yes

3. Meeting time

Two towns are 330 km apart. A car leaves Town A toward Town B at 9:00 AM at 60 km/h. Another leaves Town B toward Town A at 10:00 AM at 90 km/h. Think carefully and give only the meeting time in 12-hour format.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 667 676 12.234 s 11:48 AM Yes
Baseline 1,204 1,213 21.345 s 11:48 AM Yes

4. Coin probability

A fair coin is flipped five times. Think carefully and give only the probability of getting exactly two heads as a reduced fraction.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 439 446 8.284 s 5/16 Yes
Baseline 743 750 13.270 s 5/16 Yes

5. Python mutation trace

What does this Python code print? Think through mutation order and give only the integer.

x = [1, 2, 3, 4]
for i in range(len(x)):
    x[i] += sum(x[:i])
print(x[-1])
Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 756 761 13.067 s 15 Yes
Baseline 1,629 1,634 28.674 s 15 Yes

6. Age algebra

Six years ago Alice was twice Bob's age. Six years from now Alice will be one and a half times Bob's age. Think carefully and give only Alice's current age as an integer.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 554 559 9.648 s 30 Yes
Baseline 1,038 1,043 17.862 s 30 Yes

7. Task ordering

Four distinct tasks A, B, C, and D must be ordered. A must occur before C; B must occur before C; and B must occur before D. Think carefully and give only the number of valid complete orderings.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 2,196 2,204 43.488 s 5 Yes
Baseline 3,992 4,096 76.027 s None; token limit reached No

The baseline's true reasoning requirement for this prompt is unknown because generation was truncated at the shared 4,096-token limit.

8. Rectangle area

A rectangle has perimeter 66. Its length is 3 more than twice its width. Think carefully and give only its area as an integer.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 428 434 7.609 s 230 Yes
Baseline 692 698 12.286 s 230 Yes

9. Calendar reasoning

In a non-leap year, January 1 is a Monday. Think carefully and give only the weekday on March 1.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 438 442 8.572 s Thursday Yes
Baseline 1,117 1,121 20.997 s Thursday Yes

10. Chickens and rabbits

A farm has only chickens and rabbits: 35 heads and 100 legs total. Think carefully and give only the number of rabbits.

Model Reasoning tokens Completion tokens Time Final answer Correct
ThinkingCap 378 383 6.658 s 15 Yes
Baseline 495 500 8.919 s 15 Yes

Interpretation

On this small deterministic test, ThinkingCap preserved answer quality while substantially reducing internal reasoning. The lower wall-clock time closely tracked the lower token count, while MTP acceptance was slightly lower for ThinkingCap. That suggests the speedup came primarily from generating fewer reasoning tokens, not from more successful speculative decoding.

This is not a comprehensive quality benchmark. It uses one run per prompt, ten relatively compact reasoning problems, and temperature 0. Broader evaluation should include repeated runs, harder problems, coding tasks, long-context work, and non-deterministic sampling. The results do nevertheless support the intended behavior of ThinkingCap on this workload: similar answers with materially shorter reasoning traces.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment