Fresco 1.3
Fresco and earlier Sennoric releases alongside published results from several well-known models. Glyph 1.1 is marked historical.
Tested under pressure.
Fresco 1.3 gave a safe answer in 48.9% of intentionally hostile tests. The run stopped at 188 of 300 total planned responses, so the result remains partial.
Read the full safety reportA narrow severe-harm classifier.
An external classifier is being added for a small set of severe harmful requests, including targeted hateful abuse and sexual exploitation of children. It is designed to leave normal conversations and legitimate tool calls alone.
The classifier is planned and still needs separate validation; it is not included in the model's reported safety score.
How the tests were run
HumanEval coding test: 96.4% reported; the exact test settings and saved answers still need to be attached.
GSM8K math test: all 1,319 questions completed; Fresco 1.3 answered 1,150 correctly.
BFCL V4 tool-use test: fixed 100-task pilot, with 87 correct.
IFEval instruction-following test: 494 of 541 questions completed, but the test settings caused repetitive answers, so the result is not used as a launch headline.
Safety test: 188 of 300 planned responses completed; 48.9% is below the stated 80% internal target.