Aramis's picture
38 6

Aramis

amenur

AI & ML interests

None yet

Recent Activity

Organizations

None yet

amenur's activity

upvoted an article about 12 hours ago
upvoted an article 10 days ago
view article
Article

SmolVLM2: Bringing Video Understanding to Every Device

โ€ข 205
upvoted an article 18 days ago
upvoted an article 20 days ago
view article
Article

SigLIP 2: A better multilingual vision language encoder

โ€ข 133
upvoted an article 29 days ago
view article
Article

Open-source DeepResearch โ€“ Freeing our search agents

โ€ข 1.16k
upvoted 3 articles about 1 month ago
view article
Article

Introducing smolagents: simple agents that write actions in code.

โ€ข 867
view article
Article

Open-R1: a fully open reproduction of DeepSeek-R1

โ€ข 803
upvoted an article 2 months ago
view article
Article

Superposition in Transformers: A Novel Way of Building Mixture of Experts

By BenChaliah โ€ข
โ€ข 14
reacted to lewtun's post with ๐Ÿš€๐Ÿ”ฅ 3 months ago
view post
Post
6932
We outperform Llama 70B with Llama 3B on hard math by scaling test-time compute ๐Ÿ”ฅ

How? By combining step-wise reward models with tree search algorithms :)

We show that smol models can match or exceed the performance of their much larger siblings when given enough "time to think"

We're open sourcing the full recipe and sharing a detailed blog post.

In our blog post we cover:

๐Ÿ“ˆ Compute-optimal scaling: How we implemented DeepMind's recipe to boost the mathematical capabilities of open models at test-time.

๐ŸŽ„ Diverse Verifier Tree Search (DVTS): An unpublished extension we developed to the verifier-guided tree search technique. This simple yet effective method improves diversity and delivers better performance, particularly at large test-time compute budgets.

๐Ÿงญ Search and Learn: A lightweight toolkit for implementing search strategies with LLMs and built for speed with vLLM

Here's the links:

- Blog post: HuggingFaceH4/blogpost-scaling-test-time-compute

- Code: https://github.com/huggingface/search-and-learn

Enjoy!
  • 2 replies
ยท
upvoted an article 6 months ago
view article
Article

Llama can now see and run on your device - welcome Llama 3.2

โ€ข 184
upvoted 2 articles 6 months ago
view article
Article

Fine-tuning LLMs to 1.58bit: extreme quantization made easy

โ€ข 225
view article
Article

Scaling robotics datasets with video encoding

โ€ข 39
upvoted an article 7 months ago
view article
Article

Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth

By mlabonne โ€ข
โ€ข 293