We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
if you think students will be routinely collaborating with LLMs after graduation, #1 priority is to study rhetoric because you need to recognize BS and critique the rigor of an argument that *looks* good. the elite users are going to be like, philosophy, classics, and english majors.
if you are a PhD student in AI, remember it is in your interests to distract your advisor from how much money they could be making in industry. should be a daily priority.
Yes, some people are worse at writing than an LLM. Those people are also usually incoherent thinkers. If they learned to write better, they would get better at thinking.
Are students embarrassed by AI cheating? Like, there has always been rare but unstigmatized dishonesty (memorizing a frat’s archive of finals questions) and common but stigmatized (saying you’ve started when you definitely have not started it). AI cheating should be the most stigmatized. Is it?
ok the thing about erdos is he obviously loved collaborating with humans. he could have done a lot on his own, but math was how he chose to connect. I'm not sure he would have been very into chatbots?
ACL needs to adopt the expectation from ML conferences that workshops be exciting. If you come from NLP you know ACL workshops are mostly terminal venues for abandoned/unambitious work, but it's clear that the correct approach is to host cutting-edge WiP.
How do you make LLMs actually good at explaining new math? It's like reading a badly written reference for people who already know the subject. When I ask a question, it never matches my level. If I try to rephrase to test my understanding, it just sycophantically agrees.
my new literary award cannot be won by a commercial frontier LLM because I will require that 10% of each submission is smut
the fact that AI judges prefer sloppy AI writing makes the total death of good human-readable prose almost inevitable in scientific publishing and writing competitions. not sure what we can do about that.