Hugging Face

Enterprise

company

Verified

https://huggingface.co.

huggingface

Activity Feed

AI & ML interests

The AI community building the future.

Recent Activity

lysandre updated a dataset about 3 hours ago

huggingface/transformers-metadata

pepijn223 updated a dataset about 4 hours ago

huggingface/documentation-images

julien-c new activity about 9 hours ago

huggingface/HuggingDiscussions:[FEEDBACK] Inference Providers

View all activity

Articles

Yay! Organizations can now publish blog Articles

Jan 20

• 41

huggingface's activity

lysandre

updated a dataset about 3 hours ago

huggingface/transformers-metadata

Viewer • Updated about 3 hours ago • 1.64k • 1.22k • 23

pepijn223

updated a dataset about 4 hours ago

huggingface/documentation-images

Viewer • Updated about 4 hours ago • 52 • 2.94M • 60

julien-c

in huggingface/HuggingDiscussions about 9 hours ago

[FEEDBACK] Inference Providers

106

#49 opened 3 months ago by

julien-c

AdinaY

posted an update about 9 hours ago

Post

1734

Kimi-Audio 🚀🎧 an OPEN audio foundation model released by Moonshot AI
moonshotai/Kimi-Audio-7B-Instruct
✨ 7B
✨ 13M+ hours of pretraining data
✨ Novel hybrid input architecture
✨ Universal audio capabilities (ASR, AQA, AAC, SER, SEC/ASC, end-to-end conversation)

a-r-r-o-w

in huggingface/documentation-images about 15 hours ago

Uploading all necessary assets for VisualCloze at Dffusers.

#483 opened 3 days ago by

lzyhha

julien-c

in huggingface/HuggingDiscussions 2 days ago

[FEEDBACK] Local apps

#31 opened 11 months ago by

kramp

julien-c

in huggingface/the-no-branch-repo 2 days ago

Create README.md

#1 opened 18 days ago by

Gajduk

julien-c

in huggingface/transformers-metadata 2 days ago

Create Sporty-Video-Download,cutting app

#4 opened 16 days ago by

joerg23

julien-c

in huggingface/brand-assets 2 days ago

Request for social media icon vector

#4 opened 4 months ago by

umarbutler

julien-c

in huggingface/label-files 2 days ago

Rename lvis-id2label.json to trees.json

#12 opened 25 days ago by

HARENDRAKDAD

Upload trees.json

#10 opened 25 days ago by

HARENDRAKDAD

julien-c

posted an update 3 days ago

Post

3543

BOOOOM: Today I'm dropping TINY AGENTS

the 50 lines of code Agent in Javascript 🔥

I spent the last few weeks working on this, so I hope you will like it.

I've been diving into MCP (Model Context Protocol) to understand what the hype was all about.

It is fairly simple, but still quite powerful: MCP is a standard API to expose sets of Tools that can be hooked to LLMs.

But while doing that, came my second realization:

Once you have a MCP Client, an Agent is literally just a while loop on top of it. 🤯

➡️ read it exclusively on the official HF blog: https://huggingface.co./blog/tiny-agents

1 reply

julien-c

updated a dataset 3 days ago

huggingface/documentation-images

Viewer • Updated about 4 hours ago • 52 • 2.94M • 60

pagezyhf

updated a dataset 3 days ago

huggingface/documentation-images

Viewer • Updated about 4 hours ago • 52 • 2.94M • 60

merve

posted an update 3 days ago

Post

3249

Don't sleep on new AI at Meta Vision-Language release! 🔥

facebook/perception-encoder-67f977c9a65ca5895a7f6ba1
facebook/perception-lm-67f9783f171948c383ee7498

Meta dropped swiss army knives for vision with A2.0 license 👏
> image/video encoders for vision language modelling and spatial understanding (object detection etc) 👏
> The vision LM outperforms InternVL3 and Qwen2.5VL 👏
> They also release gigantic video and image datasets

The authors attempt to come up with single versatile vision encoder to align on diverse set of tasks.

They trained Perception Encoder (PE) Core: a new state-of-the-art family of vision encoders that can be aligned for both vision-language and spatial tasks. For zero-shot image tasks, it outperforms latest sota SigLIP2 👏

> Among fine-tuned ones, first one is PE-Spatial. It's a model to detect bounding boxes, segmentation, depth estimation and it outperforms all other models 😮

> Second one is PLM, Perception Language Model, where they combine PE-Core with Qwen2.5 LM 7B. it outperforms all other models (including InternVL3 which was trained with Qwen2.5LM too!)

The authors release the following checkpoints in sizes base, large and giant:

> 3 PE-Core checkpoints (224, 336, 448)
> 2 PE-Lang checkpoints (L, G)
> One PE-Spatial (G, 448)
> 3 PLM (1B, 3B, 8B)
> Datasets

Authors release following datasets 📑
> PE Video: Gigantic video datasete of 1M videos with 120k expert annotations ⏯️
> PLM-Video and PLM-Image: Human and auto-annotated image and video datasets on region-based tasks
> PLM-VideoBench: New video benchmark on MCQA

2 replies

sayakpaul

authored a paper 5 days ago

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning

Paper • 2504.16080 • Published 6 days ago • 15

AdinaY

posted an update 5 days ago

Post

3143

MAGI-1 🪄 the autoregressive diffusion video model, released by Sand AI

sand-ai/MAGI-1

✨ 24B with Apache 2.0
✨ Strong temporal consistency
✨ Benchmark-topping performance

1 reply

merve

posted an update 5 days ago

Post

3076

New foundation model on image and video captioning just dropped by NVIDIA AI 🔥

Describe Anything Model (DAM) is a 3B vision language model to generate detailed captions with localized references 😮

The team released the models, the dataset, a new benchmark and a demo 🤩 nvidia/describe-anything-680825bb8f5e41ff0785834c

Most of the vision LMs focus on image as a whole, lacking localized references in captions, and not taking in visual prompts (points, boxes, drawings around objects)

DAM addresses this on two levels: new vision backbone that takes in focal crops and the image itself, and a large scale dataset 👀

They generate a dataset by extending existing segmentation and referring expression generation datasets like REFCOCO, by passing in the images and classes to VLMs and generating captions.

Lastly, they also release a new benchmark again with self-supervision, they use an LLM to evaluate the detailed captions focusing on localization 👏

davanstrien

posted an update 5 days ago

Post

1811

Came across a very nice submission from @marcodsn for the reasoning datasets competition (https://huggingface.co./blog/bespokelabs/reasoning-datasets-competition).

The dataset distils reasoning chains from arXiv research papers in biology and economics. Some nice features of the dataset:

- Extracts both the logical structure AND researcher intuition from academic papers
- Adopts the persona of researchers "before experiments" to capture exploratory thinking
- Provides multi-short and single-long reasoning formats with token budgets - Shows 7.2% improvement on MMLU-Pro Economics when fine-tuning a 3B model

It's created using the Curator framework with plans to scale across more scientific domains and incorporate multi-modal reasoning with charts and mathematics.

I personally am very excited about datasets like this, which involve creativity in their creation and don't just rely on $$$ to produce a big dataset with little novelty.

Dataset can be found here: marcodsn/academic-chains (give it a like!)

meg

posted an update 6 days ago

Post

2240

New launch: See the energy use of chatbot conversations, in real time. =)
jdelavande/chat-ui-energy
Great work from @JulienDelavande !

AI & ML interests

Recent Activity

Articles

Yay! Organizations can now publish blog Articles

Team members 212

huggingface's activity

[FEEDBACK] Inference Providers

Uploading all necessary assets for VisualCloze at Dffusers.

[FEEDBACK] Local apps

Create README.md

Create Sporty-Video-Download,cutting app

Request for social media icon vector

Rename lvis-id2label.json to trees.json

Upload trees.json