Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,364 searchable experts 2,966 tracked across all sources
Clear
Showing signals surfaced by Vilém Zouhar ×

The developments commanding sustained expert attention

One card per development. Sources are clustered; expert reactions remain attributable.

Established Evaluation & benchmarks Resource 6d ago
Last Translation Benchmark tool

Last Translation Benchmark

3 experts are actively discussing the implications.

1 expert 1 community 1 sources clustered

“Today we're opening the platform to public contributors with first release planned on September 1st. Don't miss out! 🪨 Contribute here: last-translation-benchmark.vilda.net”

3 experts discussed this · 9 posts
Vilém Zouhar: There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper).
Vilém Zouhar: Multiple things made us start this effort. Typical translation benchmarks are.. ...oftentimes trivial or saturated (so they can't be used for guiding the next steps in the field) ...not evaluatable…
Vilém Zouhar: In the Last Translation Benchmark we solve both by: - collecting hard-to-translate inputs (texts, images, audios) - requiring human-readable "verification rules", which enable provable evaluation o…
Open the full discussion →