UCSF-JHU Opioid Industry Documents Archive

university

https://www.industrydocuments.ucsf.edu/opioids/

Activity Feed

AI & ML interests

None defined yet.

Recent Activity

dvilasuero authored a paper 20 days ago

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

bwingenroth updated a dataset about 1 month ago

opioidarchive/oida-qa

khawki29 updated a dataset about 1 month ago

opioidarchive/oida-qa

View all activity

opioidarchive's activity

davanstrien

posted an update 6 days ago

Post

1544

Introducing FineWeb-C 🌐🎓, a community-built dataset for improving language models in ALL languages.

Inspired by FineWeb-Edu the community is labelling the educational quality of texts for many languages.

318 annotators, 32K+ annotations, 12 languages - and growing! 🌍

data-is-better-together/fineweb-c

dvilasuero

authored a paper 20 days ago

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Paper • 2412.03304 • Published 22 days ago • 17

dvilasuero

posted an update 20 days ago

Post

2261

🌐 Announcing Global-MMLU: an improved MMLU Open dataset with evaluation coverage across 42 languages, built with Argilla and the Hugging Face community.

Global-MMLU is the result of months of work with the goal of advancing Multilingual LLM evaluation. It's been an amazing open science effort with collaborators from Cohere For AI, Mila - Quebec Artificial Intelligence Institute, EPFL, Massachusetts Institute of Technology, AI Singapore, National University of Singapore, KAIST, Instituto Superior Técnico, Carnegie Mellon University, CONICET, and University of Buenos Aires.

🏷️ +200 contributors used Argilla MMLU questions where regional, dialect, or cultural knowledge was required to answer correctly. 85% of the questions required Western-centric knowledge!

Thanks to this annotation process, the open dataset contains two subsets:

1. 🗽 Culturally Agnostic: no specific regional, cultural knowledge is required.
2. ⚖️ Culturally Sensitive: requires dialect, cultural knowledge or geographic knowledge to answer correctly.

Moreover, we provide high quality translations of 25 out of 42 languages, thanks again to the community and professional annotators leveraging Argilla on the Hub.

I hope this will ensure a better understanding of the limitations and challenges for making open AI useful for many languages.

Dataset: CohereForAI/Global-MMLU

davanstrien

posted an update 27 days ago

Post

490

Increasingly, LLMs are becoming very useful for helping scale annotation tasks, i.e. labelling and filtering. When combined with the structured generation, this can be a very scalable way of doing some pre-annotation without requiring a large team of human annotators.

However, there are quite a few cases where it still doesn't work well. This is a nice paper looking at the limitations of LLM as an annotator for Low Resource Languages: On Limitations of LLM as Annotator for Low Resource Languages (2411.17637).

Humans will still have an important role in the loop to help improve models for all languages (and domains).

davanstrien

posted an update 30 days ago

Post

2473

First dataset for the new Hugging Face Bluesky community organisation: bluesky-community/one-million-bluesky-posts 🦋

📊 1M public posts from Bluesky's firehose API
🔍 Includes text, metadata, and language predictions
🔬 Perfect to experiment with using ML for Bluesky 🤗

Excited to see people build more open tools for a more open social media platform!

davanstrien

posted an update about 1 month ago

Post

1349

The Bluesky AT Protocol unlocks exciting possibilities:
- Building custom feeds using ML
- Creating dashboards for data exploration
- Developing custom models for Bluesky
To gather Bluesky resources on the Hub, I've created a community org: https://huggingface.co/bluesky-community

My first rather modest contribution is a dashboard that shows the number of posts every second. Drinking straight from the firehose API 🚰

bluesky-community/bluesky-posts-over-time

1 reply

bwingenroth

updated a dataset about 1 month ago

opioidarchive/oida-qa

Viewer • Updated Nov 22 • 400k • 437

khawki29

updated a dataset about 1 month ago

opioidarchive/oida-qa

Viewer • Updated Nov 22 • 400k • 437

shawnricecake

updated a dataset about 1 month ago

opioidarchive/oida-qa

Viewer • Updated Nov 22 • 400k • 437

davanstrien

posted an update about 1 month ago

Post

1306

huggingface.co/DIBT is dead!

Long live https://huggingface.co/data-is-better-together!

We're working on some very cool projects so we're doing a bit of tidying of the Data is Better Together Hub org 🤓

dvilasuero

posted an update about 1 month ago

Post

1106

@Jesse-marqo and the Marqo team are killing it on the Hub: top embedding models and datasets!

Here's how to start using their new evaluation dataset for curation and labelling:

1. Deploy Argilla on Spaces: https://huggingface.co/new-space?template=argilla%2Fargilla-template-space
2. Load Marqo/amazon-products-eval with the UI wizard.
3. Start curating!

JoshuaGu

authored a paper about 2 months ago

Personalization of Large Language Models: A Survey

Paper • 2411.00027 • Published Oct 29 • 31

dvilasuero

posted an update about 2 months ago

Post

683

Build datasets for AI on the Hugging Face Hub—10x easier than ever!

Today, I'm excited to share our biggest feature since we joined Hugging Face.

Here’s how it works:

1. Pick a dataset—upload your own or choose from 240K open datasets.
2. Paste the Hub dataset ID into Argilla and set up your labeling interface.
3. Share the URL with your team or the whole community!

And the best part? It’s:
- No code – no Python needed
- Integrated – all within the Hub
- Scalable – from solo labeling to 100s of contributors

I am incredibly proud of the team for shipping this after weeks of work and many quick iterations.

Let's make this sentence obsolete: "Everyone wants to do the model work, not the data work."

Read, share, and like the HF blog post:
https://huggingface.co/blog/argilla-ui-hub

JoshuaGu

updated a dataset about 2 months ago

opioidarchive/oida-qa

Viewer • Updated Nov 22 • 400k • 437

davanstrien

updated a Space about 2 months ago

Sleeping

🚀

Argilla Progress

davanstrien

posted an update about 2 months ago

Post

2527

Excited to see my weird davanstrien/ufo-ColPali dataset featured in a video by @sabrinaesaquino !

The video covers using ColPali with Binary Quantization in Qdant to accelerate retrieval. 2x speed up with no performance drop in results 🛸

Video: https://youtu.be/_A90A-grwIc?si=oB3JAhJG8VQUZGLz
Blog post: https://danielvanstrien.xyz/posts/post-with-code/colpali-qdrant/2024-10-02_using_colpali_with_qdrant.html

2 replies

dvilasuero

posted an update 2 months ago

Post

989

Big news! You can now build strong ML models without days of human labelling

You simply:
- Define your dataset, including annotation guidelines, labels and fields
- Optionally label some records manually.
- Use an LLM to auto label your data with a human (you? your team?) in the loop!

Get started with this blog post:
https://huggingface.co/blog/sdiazlor/custom-text-classifier-ai-human-feedback

khawki29

updated a Space 3 months ago

Running

📈

README

davanstrien

updated a Space 3 months ago

Running on CPU Upgrade

✍

Argilla

davanstrien

posted an update 3 months ago

Post

1245

ColPali is an exciting new approach to multimodal document retrieval, but some doubt its practical use with existing vector DBs.

It turns out it's super easy to use Qdrant to index and search ColPali embeddings efficiently.

Blog post here: https://danielvanstrien.xyz/posts/post-with-code/colpali-qdrant/2024-10-02_using_colpali_with_qdrant.html

Very silly demo: davanstrien/ufo-ColPali-Search

AI & ML interests

Recent Activity

Team members 6

opioidarchive's activity

Argilla Progress

README

Argilla