donut-earth (Donut Earthers 🍩)

posted an update 4 days ago

Post

1639

~75% on the challenging GPQA with only 40M parameters 🔥🥳

GREAT ACHIEVEMENT ! Or is it ?

This new Work, "Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation", take out the mystery about many models i personally suspected their results. Speacially on leaderboards other than the english one, Like the Open Arabic LLM Leaderbaord OALL/Open-Arabic-LLM-Leaderboard.

The authors of this work, first started by training a model on the GPQA data, which, unsurprisingly, led to the model achieving 100% performance.

Afterward, they trained what they referred to as a 'legitimate' model on legitimate data (MedMCQA). However, they introduced a distillation loss from the earlier, 'cheated' model.

What they discovered was fascinating: the knowledge of GPQA leaked through this distillation loss, even though the legitimate model was never explicitly trained on GPQA during this stage.

This raises important questions about the careful use of distillation in model training, especially when the training data is opaque. As they demonstrated, it’s apparently possible to (intentionally or unintentionally) leak test data through this method.

Find out more: Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation (2412.15255)

1 reply

·

takarajordan

posted an update 16 days ago

Post

1112

I made an RSS feed for HuggingFace Daily Papers!! 🤗

Just Subscribe here: https://papers.takara.ai/api/feed

It updates every 24 hours, completely written as a serverless go script with a Redis cache (to avoid hitting HF all the time).

I'm open sourcing the code, you can check out my repo and deploy it on Vercel extremely easily!
https://github.com/404missinglink/HF-Daily-Papers-Feeds

thanks to @John6666 @p3nGu1nZz for your early support

alielfilali01

posted an update 21 days ago

Post

3377

Unpopular opinion: Open Source takes courage to do !

Not everyone is brave enough to release what they have done (the way they've done it) to the wild to be judged !
It really requires a high level of "knowing wth are you doing" ! It's kind of a super power !

Cheers to the heroes here who see this!

3 replies

·

takarajordan

posted an update 23 days ago

Post

2251

I'm super excited to release my first open-source text dataset:

WorldScenario 20K is a novel dataset of 20,000 synthetically generated multi-stakeholder scenarios designed to simulate real-world decision-making processes. Each scenario explores a unique environmental, societal, or economic issue.

I used the brand new meta-llama/Llama-3.3-70B-Instruct model to generate this dataset and I put the dataset through some post processing to clean and evaluate the dataset for diversity.

I'd appreciate some feedback and thoughts on my new release! Thanks!

takarajordan/WorldScenario_20K

8 replies

·

alielfilali01

posted an update 25 days ago

Post

1504

Apparently i forgot to put this here !

Well, this is a bit late but consider given our recent blog a read if you are interested in Evaluation.

You don't have to be into Arabic NLP in order to read it, the main contribution we are introducing is a new evaluation measure for NLG. We made the fisrt application of this measure on Arabic for now and we will be working with colleagues from the community to expand it to other languages.

Blog:
Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard
https://huggingface.co/blog/leaderboard-3c3h-aragen

Space:
inceptionai/AraGen-Leaderboard

Give it a read and let me know your thoughts 🤗

p3nGu1nZz

in donut-earth/donut-AE about 1 month ago

enhance-setup-script

4

#1 opened about 1 month ago by

p3nGu1nZz

cappuch

in donut-earth/donut-AE about 1 month ago

enhance-setup-script

4

#1 opened about 1 month ago by

p3nGu1nZz

updated a model about 1 month ago

donut-earth/donut-AE

Updated Dec 1, 2024 • 3

cappuch

updated a Space about 1 month ago

Running

1

🚀

README

cappuch

updated a model about 1 month ago

donut-earth/donut-AE

Updated Dec 1, 2024 • 3

takarajordan

posted an update about 1 month ago

Post

1208

I'm not sure why I haven't done this already!

I just made a space to count and visualize tokens for Diffusion models, no more guesswork! It's super fast too.

Check it out here and try out your prompts: takarajordan/DiffusionTokenizer

Uses these tokenizers below:
openai/clip-vit-large-patch14
google/t5-v1_1-xxl

cappuch

posted an update about 1 month ago

Post

1581

p104-100s are beasts. 8 gigs of VRAM, 12 tok/s on qwen 14b at q4, and 18 tok/s on 7b at q6. best thing - 20 euros each.

https://furry.engineer/@cappuch/113500349547803802

3 replies

·

takarajordan

updated a dataset about 1 month ago

donut-earth/proof

Viewer • Updated Nov 21, 2024 • 42 • 922 • 5

takarajordan

in donut-earth/proof about 1 month ago

Another fake image

#4 opened about 1 month ago by

qkasriel

Fake image in the dataset.

1

#1 opened about 1 month ago by

qkasriel

not-lain

updated a Space about 1 month ago

Running

1

🚀

README

takarajordan

posted an update about 1 month ago

Post

1121

First post here goes!

takarajordan/CineDiffusion

Super excited to announce CineDiffusion🎥, it creates images up to 4.2 Megapixels in Cinematic ultrawide formats like:
- 2.39:1 (Modern Widescreen)
- 2.76:1 (Ultra Panavision 70)
- 3.00:1 (Experimental Ultra-wide)
- 4.00:1 (Polyvision)
- 2.55:1 (CinemaScope)
- 2.20:1 (Todd-AO)

More to come soon!!

Thanks to @John6666 and @Resoldjew for your early support <3

And thanks to the team at ShuttleAI for their brand new Shuttle-3 model, what an amazing job.

shuttleai/shuttle-3-diffusion

not-lain

posted an update about 2 months ago

Post

1990

ever wondered how you can make an API call to a visual-question-answering model without sending an image url 👀

you can do that by converting your local image to base64 and sending it to the API.

recently I made some changes to my library "loadimg" that allows you to make converting images to base64 a breeze.
🔗 https://github.com/not-lain/loadimg

API request example 🛠️:

from loadimg import load_img
from huggingface_hub import InferenceClient

# or load a local image
my_b64_img = load_img(imgPath_url_pillow_or_numpy ,output_type="base64" ) 

client = InferenceClient(api_key="hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx")

messages = [
	{
		"role": "user",
		"content": [
			{
				"type": "text",
				"text": "Describe this image in one sentence."
			},
			{
				"type": "image_url",
				"image_url": {
					"url": my_b64_img # base64 allows using images without uploading them to the web
				}
			}
		]
	}
]

stream = client.chat.completions.create(
    model="meta-llama/Llama-3.2-11B-Vision-Instruct", 
	messages=messages, 
	max_tokens=500,
	stream=True
)

for chunk in stream:
    print(chunk.choices[0].delta.content, end="")

alielfilali01

posted an update about 2 months ago

Post

2183

Unpopular opinion : o1-preview is more stupid than 4o and Qwen2.5-72B-Instruct in extremely underrated !

2 replies

·

alielfilali01

posted an update 2 months ago

Post

1701

I feel like this incredible resource hasn't gotten the attention it deserves in the community!

@clefourrier and generally the HuggingFace evaluation team put together a fantastic guidebook covering a lot about 𝗘𝗩𝗔𝗟𝗨𝗔𝗧𝗜𝗢𝗡 from basics to advanced tips.

link : https://github.com/huggingface/evaluation-guidebook

I haven’t finished it yet, but i'am enjoying every piece of it so far. Huge thanks @clefourrier and the team for this invaluable resource !

3 replies

·

Donut Earthers 🍩

AI & ML interests

Recent Activity

donut-earth's activity

enhance-setup-script

enhance-setup-script

donut-earth/donut-AE

README

donut-earth/donut-AE

donut-earth/proof

Another fake image

Fake image in the dataset.

README

AI & ML interests

Recent Activity

Team members 8

donut-earth's activity

enhance-setup-script

enhance-setup-script

README

Another fake image

Fake image in the dataset.

README