Open source the post training methodology or dataset and is V3 also in the works?

#10
by jtvino - opened

Any plans on open sourcing the methodology for post training or releasing the dataset so it can be verified?

We finetuned the MAI-DS-R1 model on a carefully curated set of ~350K blocked topics examples using various strategies:  

Collecting and filtering query keywords 
Converting the keywords into multiple questions 
Translating the questions into various languages 
Bootstrapping answers and respective Chain of Thought (CoT) for these questions using DeepSeek R1 and internal models. 

Also is MAI-DS-V3-0324 version of this in the works

Your need to confirm your account before you can post a new comment.

Sign up or log in to comment