The Data Labeling Marketplace | Where AI Builders and AI Trainers Connect to Build the Future | Find, hire, & securely pay data labelers for any annotation toolopentrain.ai Seattle, WashingtonJoined September 2023
@Suhail Here to help! opentrain.ai
Post your data labeling project and connect with freelancers and/or data labeling agencies.
We also offer a fully-managed solution, where we build your own annotation team and manage the whole process.
DMs are also open.
At this point I feel like we understand pretty well what's going on with LLMs:
- Outputs are roughly equivalent to kernel smoothing over positional embeddings (arxiv.org/pdf/1908.11775…)
- The learned computation model is *probably* bounded by RASP-L (arxiv.org/pdf/2310.16028…)
- LLMs learn structure primarily from human generated content (text, images) which is far more structured and predictable than the universe.
- LLaMa3 shows us that the higher quality the annotations on the human generated content, the better the LLMs do (10million messages is a lot!)
- Multi-turn labeling is currently very expensive and so likely driving the costs of the models.
- Right now we're likely bottlenecked not on CPU, or size of data, but number and quality of annotations.
So tl;dr. Great at predicting what a human would do or say by averaging in distribution data in the corpus. No emergent generality. Currently bottlenecked by high quality annotated data.
@hausdorff_space did I miss anything?
The perfect quote to describe LLMs can be found in a 1946 Jean Cocteau movie -- "Réfléchissez pour moi, je réfléchirai pour vous" (think for me, I will reflect for you).
What you get from the model is always a reflection of the training data you put in -- itself a by-product of human thinking.
NEW FEATURE:
AI Live Chat Interviews: Using GPT-4 to interview, score, and screen job applicants from data labelers/AI trainers.
One-on-one interviews, at scale!
How It Works:
#GPT4
With the power of LLMs, we can now do all of this for you behind-the-scenes and then present an AI Qualification Score.
From there, you can get a clear view of who are the most qualified candidates.
Traditionally, selecting who to hire required significant time for:
-Viewing applicants' resume
-Reading cover letters
-Messaging follow up questions
-THEN judge if they are a good fit for your specific job
@ylecun We're working on this. Hire & pay data labelers/contributors directly, with only a small transaction fee. But if the resulting dataset will be open sourced, ZERO FEES.
OpenTrain.ai
Agreed! But some datasets are just too large/require too much time. This is when many turn to crowdsourcing. The issue with crowdsourcing is it's largely a "black box". You don't get any feedback on issues with the dataset or anomalies. This is becoming even more important as people start building more advanced models. Advanced models require labelers with domain-expertise that is not available through crowdsourcing.
Here's the data labeling process we'd recommend:
1) Data scientist/engineer do the initial labeling (10-20 hours recommended)
2) Source & hire INDIVIDUAL data labelers with expertise in your data's domain (on OpenTrain.ai of course 😉)
3) Work closely with the data labelers (daily sync up, open communication channels, etc)
This allows for deep understanding of training data, while still being scalable for large datasets.
A big reason for this is the cost of hiring data talent. The big data labeling platforms are prohibitively expenses and openly have a 70% margin (and pay their data labelers poorly). We're working on making it easier & more cost-effective for the open source community to get custom data by allowing people to post a dataset need > Get proposals from freelancers (many of whom also work for the large data providers) > Pay them directly with only a small transaction fee for completed milestones. OpenTrain.ai ...We also wave the transaction fee if the resulting data set will be open sourced
We're working on a solution for this. There are a lot of great tools, like Label Studio, but finding specialized talent to do the actual annotation work can be difficult.
With OpenTrain, you post your job and select your labeling software > workers submit proposals > you hire and simply add the workers to your labeling software account.
You then pay the talent directly in escrow milestones. Also, if the specialized talent isn't on our platform yet, we recruit them and bring them onboard for your job.
We only charge a small % on successful milestone payouts...vs the big data labeling companies who charge a massive markup and long term contracts.
We support 20+ data labeling tools and hiring in 110+ countries.
Grok-1 by @xai utilized "AI Tutors", or human subject matter experts to create custom training data & provide RLHF.
We're making this easy for anybody to do this with OpenTrain.ai: The data labeling marketplace to find, hire, & pay training data experts for ANY data labeling software.
Simply setup the data workflows on your data labeling tool of choice, & use OpenTrain to find the subject matter experts to work from that tool.
Like UpWork, but for "AI Tutoring"
Post jobs & collect applicants 100% free!
#rlhf#llm#datalabeling#trainingdata
🔥Excited to introduce LMSYS-Chat-1M, a large-scale dataset of 1M real-world conversations with 25 cutting-edge LLMs!
This dataset, collected from chat.lmsys.org, offers insights into user interactions with LLMs and intriguing use cases.
Link: huggingface.co/datasets/lmsys…
An in-depth look at RLHF by @natolambert from @huggingface. The need for high-quality, task-specific data in RLHF is crucial. With OpenTrainAI, you can find, hire, & pay the human experts essential for responsible and effective RLHF. Post your job today! #RLHF#MachineLearning
Reinforcement Learning from Human Feedback (RLHF) is gaining traction. This field aims to make AI more responsible by including human values and preferences.
In this video, @natolambert, a research scientist and RLHF team lead at @huggingface explores its inner workings,
44 Followers 537 FollowingFounder, Cookiy Labs (@CookiyLabs): a Human Data Layer for research & AI. ex-Meta, TikTok, Alibaba, Tencent. CS PhD dropout @UMich.
21 Followers 135 FollowingResearch & M&E | Data Analysis 📊 | Excel | Power BI | MEAL | Building practical research, M&E & data projects | Open to NGO & development roles 🌍
562 Followers 3K FollowingJapan-born. Remote × AI × Japanese.
Quietly working on AI training & Japanese-language tasks.
Studying neuroscience & psychology.
DnB, Survival&craft&build game
0 Followers 23 FollowingPassionate about connecting with people worldwide, building meaningful friendships, and creating a strong global community that helps make the world a better.
1 Followers 11 FollowingAI Video Creator | Cinematic AI Visuals | Short-form videos/ads | Image to video | AI Film making | Content Creation |Available for freelance work
266 Followers 5K FollowingIf I discover a form of art I love, I'll retweet and express my admiration with ❤️. You might catch a glimpse of a few pieces from my own illustrations here.
43 Followers 204 FollowingVisual data provider with decentralized real-time data nodes creating digital twins of the world, measuring traffic, demand, cities, and urban intelligence.
156 Followers 319 FollowingWeb3 Community Alchemist | Community Mod |Defi Analyst | Content Marketer for Crypto Brands | Building Engaged Tribes & Alpha Content | Your growth partner💪
2K Followers 604 FollowingCo-founder and CEO @SafetyWing (YC18), building a global social safety net. Tweeting about remote work, philosophy, and the future.
4K Followers 14 FollowingUnderstanding Intelligence.
Measurement. Explanation. Application. That's how we're tackling AI interpretability: the greatest scientific problem of our age.
95K Followers 155 FollowingBuilding beautiful things like Mojo🔥 and MAX @Modular, lifting the world of production AI/ML software into a new phase of innovation. We’re hiring! 🚀🧠
22K Followers 57 FollowingWeb infrastructure for AI to search, extract, monitor, and reason over the world's information.
Start here → https://t.co/0Enegj3dhv
24K Followers 742 FollowingFounder and CEO @Zipline. Previously, built DNA/RNA computers and climbing walls. Harvard grad. Ex-professional rock climber.
35K Followers 3K FollowingFounder and CEO of @MentavaInc. Building early literacy software for preschoolers. Help me bring back excellence in education. Father of 4.
1K Followers 817 FollowingRun my company w/ 21 AI agents — each with a charter, ledger & sleep. Sharing what actually works, real numbers. Founder @vozoai, 3× Product Hunt · ex-Googler
35 Followers 137 FollowingNeed data for your AI? Sure, let random clickers label your brain scans. What could go wrong?🤷
I talk about why 99% of the industry gets AI datasets wrong
8K Followers 3K FollowingFounder and Director of the @MetascienceObs.
"Speak your mind, even if your voice shakes".
Leave anonymous feedback here: https://t.co/LQ5eZWwDst
56K Followers 1K FollowingWe build fresher maps for humanity.
Join a decentralized global community of mappers. Earn rewards.
(We don't do airdrops, etc. )