90trust / 100

Nonobench

by nonobench.com in Developer tools

MCP serverPassing, checked 5 h ago

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

https://www.nonobench.com/mcp

Last 30 days

All checks passedSome failedAll failedNot checked
Uptime
100%
Response time
221 ms typical, 221 ms slowest 5%
Last check
5 h ago
Next check
in 1 h

How to call it

Add it to any MCP client that supports remote servers.

{
  "mcpServers": {
    "nonobench": {
      "type": "http",
      "url": "https://www.nonobench.com/mcp"
    }
  }
}

11 tools

  • get_leaderboard

    Models ranked by accuracy; defaults to all effort levels for compatibility.

  • list_providers

    Provider ids, names, families and variant counts.

  • list_families

    Model families, available efforts and best variants.

  • compare_models

    Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.

  • get_model_results

    Accuracy, cost, latency and token use for one model, broken down by grid size.

  • list_puzzles

    The benchmark puzzles with their ids and row/column clues.

  • get_puzzle

    One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.

  • check_solution

    Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.

  • get_puzzle_results

    Per-model outcomes for one puzzle. Answers are omitted unless requested.

  • get_model_puzzles

    Which puzzles one model solved, missed, timed out on, or has not run.

  • list_runs

    Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.

Security scan

  • No findings. We scan names, descriptions and tool definitions for hidden instructions and other prompt-injection patterns.

Recent checks

WhenResultHTTPTime
5 h agoPassed200221 ms