Sennoric model launch · Draft

Fresco 1.3

A new foundation.

The strongest coding and reasoning model in the Solan family yet: 96.4% on HumanEval, 87.19% on GSM8K, and 87 out of 100 on a BFCL V4 tool-use test.

96.4%HumanEval · coding
87.19%GSM8K · math
87/100BFCL V4 · tool use
Benchmark results

Fresco 1.3

Fresco and earlier Sennoric releases alongside published results from several well-known models. Glyph 1.1 is marked historical.

HumanEval · coding
Percentage of coding problems solved · higher is better
Published results
Fresco 1.3
96.4%
GPT-4o
90.6%
Llama 3.1 405B
89.0%
GPT-4o mini
86.2%
Phi-4
82.6%
Llama 3.1 70B
80.5%
Qwen 2.5 72B
80.4%
Fresco 1.2
80.0%
Llama 3.3 70B
78.9%
Llama 3.1 8B
72.6%
Fresco 1.2.1
50.0%
Glyph 1.1 historical
3.0%
0%25%50%75%100%
External scores: Microsoft Phi-4 model card and Meta Llama 3.1 model card. Published test settings differ between providers, so this is a reported-score comparison rather than one controlled run.
GSM8K · math
Percentage of math questions answered correctly · higher is better
Published results
Llama 3.1 405B
96.8%
Llama 3.1 70B
95.1%
Llama 3 70B
93.0%
Fresco 1.3
87.19%
Llama 3.1 8B
84.5%
Llama 3 8B
80.6%
Fresco 1.2.1
12.0%
Fresco 1.2
5.0%
Glyph 1.1 historical
Not reported
0%25%50%75%100%
External scores: official Meta Llama 3.1 model card. Meta used its published eight-example reasoning setup; Fresco used its own complete 1,319-question run, so settings are not identical.
Safety

Tested under pressure.

Fresco 1.3 gave a safe answer in 48.9% of intentionally hostile tests. The run stopped at 188 of 300 total planned responses, so the result remains partial.

Read the full safety report
Deployment layer

A narrow severe-harm classifier.

An external classifier is being added for a small set of severe harmful requests, including targeted hateful abuse and sexual exploitation of children. It is designed to leave normal conversations and legitimate tool calls alone.

The classifier is planned and still needs separate validation; it is not included in the model's reported safety score.

How the tests were run

HumanEval coding test: 96.4% reported; the exact test settings and saved answers still need to be attached.

GSM8K math test: all 1,319 questions completed; Fresco 1.3 answered 1,150 correctly.

BFCL V4 tool-use test: fixed 100-task pilot, with 87 correct.

IFEval instruction-following test: 494 of 541 questions completed, but the test settings caused repetitive answers, so the result is not used as a launch headline.

Safety test: 188 of 300 planned responses completed; 48.9% is below the stated 80% internal target.