Catherine Arnett

catherinearnett

AI & ML interests

multilingual NLP, tokenization

Organizations

Blog-explorers's profile picture Language and Cognition Lab (UCSD)'s profile picture PleIAs's profile picture

catherinearnett's activity

published an article 5 months ago
published an article 6 months ago
view article
Article

Releasing the largest multilingual open pretraining dataset

By Pclanglais and 2 others
101
published an article 6 months ago
published an article 7 months ago