Open source the post training methodology or dataset and is V3 also in the works?
#10
by
jtvino
- opened
Any plans on open sourcing the methodology for post training or releasing the dataset so it can be verified?
We finetuned the MAI-DS-R1 model on a carefully curated set of ~350K blocked topics examples using various strategies:
Collecting and filtering query keywords
Converting the keywords into multiple questions
Translating the questions into various languages
Bootstrapping answers and respective Chain of Thought (CoT) for these questions using DeepSeek R1 and internal models.
Also is MAI-DS-V3-0324 version of this in the works