LLM code generation dramatically speeds up development until it doesn’t. If you’ve worked with AI-generated code, you’ve likely seen:
* Bugs are buried in AI spaghetti code that takes longer to fix than it does to write it yourself.
* Code that does something completely different than what you wanted
This post shows you how to prevent that. You’ll learn to specify requirements, so LLMs generate exactly what you need. And avoid wading through thousands of lines of code to verify correctness.
startdataengineering.com/post/using-llm…#dataengineering#llm#codegeneration
AI makes some reasonable guesses, but always verify the results.
I’m working on an article to streamline pipeline creation. We still need humans in the loop.
The image below recommends an ID-based delete-insert for a fact table, which will take forever to run if your data is large.
Add these to your skills.md file.
#dataengineering#llm
“I recommend this to everyone on linkedin asking "what to study" for data engineering. because the topics are so fundamental and work with literally every system (databricks, snowflake, redshift, bigquery, etc)”
- One of the comments I got about my Design Patterns workshop
Check it out here
youtu.be/AXtcWgvCUtw?t=…#dataengineering#pipelinedesign
Working on a project that teaches you how to build data pipelines following standard design practices and speeding up development with agent harness & skills!
#dataengineering#dataproject
There are a lot of data projects available on the web. While these projects are great, starting from scratch to build your data project can be challenging.
If you are wondering how to go from an idea to a production-ready data pipeline
This post is for you! In it, we will go over how to build a data project step-by-step from scratch.
By the end of this post, you will be able to quickly create data projects for any use case and see how the different parts of data systems work together.
startdataengineering.com/post/de-proj-s…#dataengineering#datapipeline
I’ve spent way too much time thinking about my brand colors.
Ultimately, I decided to go with what would make my readers' lives easier.
A bright screen makes it hard to concentrate for long intervals.
So I switched my website to use the Gruvbox theme.
Does it help?
Check it out at startdataengineering.com
In the past few years, data modeling has gone from “well-planned data model” to “dump everything in S3 and get some data for the end-user”.
How do you ensure data correctness when data is expected immediately, and there isn't enough time to get the model right?
For solutions, read this
startdataengineering.com/post/deliver-d…#dataengineering#datamodel
Updating my “Python Essentials for Data Engineers” post and came across this amazing data quality tool.
pointblank from posit.
Easy to set up and use, as any tool should be. No need to wrestle with 100s of configs/settings.
#dataengineering#python#dataquality
One of the biggest causes of data completeness issues is not knowing whether your input datasets are complete.
If you schedule pipelines guess when your upstream data may be ready, its hope driven.
Learn how to use data asset-based triggers & start your pipelines when the inputs are ready.
startdataengineering.com/post/airflow-t…#dataengineering#datapipeline#apacheairflow
Which one should I use; Gruvbox or standard?
Gruvbox is easier on the eyes
But the standard theme (Bootswatch, Flatly) is familiar
Let me know in the poll below 👇
#ux#readability
There are many data projects available online. While these projects are great, starting from scratch to build your data project can be challenging.
If you are
* Wondering how to go from an idea to a production-ready data pipeline
* Feeling overwhelmed by how all the parts of a data system fit together
* Unsure whether the pipelines you build are up to industry-standard
If so, this post is for you! In it, we will go over how to build a data project from scratch, step by step.
By the end of this post, you will be able to quickly create data projects for any use case and see how the different parts of data systems work together.
startdataengineering.com/post/de-proj-s…#dataengineering#dataprojects
Using AI for development? Go read @htmx_org ‘ s post.
htmx.org/essays/working…
Take notes.
My experience closely resembles his when building data stuff.
1. AI is good at debugging (only if we have sufficient logs & trace).
2. AI is great at creating tests (although I usually have to specify additional test cases).
3. AI will “sometimes” lead you down a complexity trap when implementing new features.
When designing new features/interfaces, I always
* Create my own plan with a clear objective and pseudocode
* Then review it with AI and iterate.
I found this to be much more time efficient and less burnout-y
#dataengineering#aidevelopment
When choosing a tool/framework, always pick the one with the simplest implementation.
But that could be extended with libraries, Plugins, Ecosystem, etc.
Here are some of my favorite tools
1. DuckDB: SQL, enhanced with extensions
2. Python: English-like syntax, enhanced with libraries
3. Quarto: Markdown-like syntax, enhanced with extensions
4. nvim: vim syntax, enhanced with extensions
The fancier/esoteric the tool, the harder it is to get productive quickly.
What are some of your favorite tools?
#tools#dataengineering
This is part of my free data design pattern workshop (next Sat), where I cover this with exercises & arguments.
Come join me, sign up here to get notified: startdataengineering.com/newsletter.
90% of all facts and dimensions follow a similar design pattern, with some differences in business logic.
This messy middle (fact & dims) is where DEs can lay the foundations for reliable data.
If you mess this up, your team will always face issues/questions/mismatches, etc.
Here is a V1 of this design
Whenever I design a workshop, I use @robfitz ’s “The Workshop Survival Guide”.
I see how great user design increases engagement, recommendations, and leads to higher ratings.
Here is a very rough version of a workshop that I am working on (coming soon).
Focus on
1. Takeaways for the audience. Not just watch and forget.
2. Increase play. Interweave code exercises and visuals to help understanding.
#dataengineering#userexperience
13K Followers 3K Following#MicrosofFabric user advocate, interests in Small Data & Self Service #Microsoftemployee since Dec 2023 , but my tweets are my own
6K Followers 2K FollowingCrafting data engineering+ stories. Educator at @sspdatahq & https://t.co/7r8pihWPQz.
Dad, Technical Author, Data Engineer. Obsidian & Neovim. Learning for Life.
5K Followers 4K Followingdata director @workingfamilies prev: @sunrisemvmt. waffle house ambassador & David Byrne fan account. budding organizer. she/her, born n raised on a holler
21 Followers 127 FollowingPaste any sheet → ATI+ detects column types, creates the Azure table & loads your data. No scripting. No setup. Just paste and go. Free on Microsoft Store
2 Followers 80 FollowingWorking as a Data Engineer. Skilled in PySpark, Python, SQL, Azure Databricks, ADF, Azure Storage, Data Modelling, Data Warehousing, RAG, AI Agent.
187 Followers 4K FollowingContent creator , developer and advocate for Linux and FOSS
صانع محتوى و مطور ، انشر عن البرمجيات الحرة و مفتوحة المصدر ، لينكس و صناعة البرمجيات
13K Followers 3K Following#MicrosofFabric user advocate, interests in Small Data & Self Service #Microsoftemployee since Dec 2023 , but my tweets are my own
6K Followers 2K FollowingCrafting data engineering+ stories. Educator at @sspdatahq & https://t.co/7r8pihWPQz.
Dad, Technical Author, Data Engineer. Obsidian & Neovim. Learning for Life.
56K Followers 90 FollowingCreator of @elixirlang. Chief Adoption Officer at @dashbit, where we build https://t.co/FK8F4URbVG and https://t.co/xncEVrvWml.
137K Followers 1K FollowingNYTimes bestselling author of STEAL LIKE AN ARTIST and other books. Subscribe to my weekly newsletter & log off this hellsite: https://t.co/dQIzmsUdAY
286K Followers 9K FollowingAuthor, explorer, xenophile, programmer, netizen, conversationalist. Former musician and entrepreneur. Everything is at https://t.co/fuY6AJRuz5
17K Followers 636 Following14 years running little businesses and 3 books about my learnings along the way. Tweets about the career path of entrepreneurship & the business of indie books.
2K Followers 3K FollowingI like to make simple helpful apps. Learn how to customize your Shopify store without coding knowledge : https://t.co/vZI65P5RF4
John 8:7
2K Followers 69 FollowingSoftware Engineer. Opinions and views shared here are my own. Most posts are jokes.
I post about things related to data and tennis.
15K Followers 720 Followingone of the first ~8,000 people on this site, somehow. left. now i'm back and working on @jfdibot to help me run my community, businesses, and more. ai mechanic.
190K Followers 423 FollowingHelping people build wealth since 2017. Author of Just Keep Buying (https://t.co/8gu4qZ7MWy) & The Wealth Ladder (https://t.co/3lGb0qPuin)