Apache Spark
69%21 posts analyzed over the last 12 weeks
Over the last 12 weeks
Average across all posts
This month vs. previous
21 posts analyzed over the last 12 weeks
Over the last 12 weeks
Average across all posts
This month vs. previous
Growth slope (6 weeks)
Best week: 20 juil. (560 avg. likes)
Matei Zaharia
CTO & Cofounder at Databricks, CS Professor at Berkeley
Lots of new goodies in Apache Spark 4.2: Metric Views, GEOGRAPHY and GEOMETRY types, AutoCDC in Spark Declarative Pipelines, and more, from 1900 total commits. Available today in Databricks: https://lnkd.in/gmiaA73G
Ajay Kadiyala
Senior Data Engineer @ ENBD | Azure • Databricks • PySpark • SQL | I simplify Data Engineering, interviews & real-world project learning for 200K+ professionals
“Should I learn Scala or PySpark for Data Engineering?” I get this question repeatedly. My answer is For most Data Engineers today, start with PySpark. But that does not mean Scala is outdated or that Scala is always …
Ajay Kadiyala
Senior Data Engineer @ ENBD | Azure • Databricks • PySpark • SQL | I simplify Data Engineering, interviews & real-world project learning for 200K+ professionals
𝗔𝘇𝘂𝗿𝗲 𝗶𝘀 𝘂𝘀𝗲𝗱 𝗯𝘆 𝟵𝟱% 𝗼𝗳 𝗙𝗼𝗿𝘁𝘂𝗻𝗲 𝟱𝟬𝟬 𝗰𝗼𝗺𝗽𝗮𝗻𝗶𝗲𝘀. 𝗧𝗵𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 𝘄𝗵𝗼 𝘄𝗼𝗿𝗸 𝘄𝗶𝘁𝗵 𝗶𝘁 𝗲𝗮𝗿𝗻 𝘀𝗼𝗺𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗵𝗶𝗴𝗵𝗲𝘀𝘁 𝘀𝗮𝗹𝗮𝗿𝗶𝗲𝘀 𝗶𝗻 𝘁𝗲𝗰𝗵. Her…
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
SQL vs PySpark: Stop choosing one. Learn how they work together. One of the biggest mistakes I see aspiring Data Engineers make is treating SQL and PySpark as competing skills. They're not. SQL teaches you how to thin…
Abhishek Agrawal
Senior Azure Data Engineer | Azure Databricks | Microsoft Fabric | Azure Data Factory | PySpark | Spark | Azure Synapse | SQL | Data Platform | Open to Germany 🇩🇪 & UAE 🇦🇪
🚀 The Ultimate PySpark Cheatsheet is Here! PySpark is one of the most essential skills for every Data Engineer—but remembering every function, optimization technique, and best practice isn't easy. That's why I created…
Serigne Faye
Data engineer | BI Engineer | #Google cloud innovators | Gcp certified
Spark ≠ PySpark… mais Spark + Python = PySpark 👉 Apache Spark : - Un moteur de calcul distribué open source. - Spécialement conçu pour le Big Data. - Ultra rapide et très flexible. - Écrit en Scala, mais prend en …
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
Most data engineers can write a df.write(), but very few can tell you exactly what happens under the hood . : : : If you want to move from writing basic queries to building production-grade, optimized data pipelines, yo…
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
Do you use PySpark... but still don't understand what happens after you click Run? If terms like these still confuse you: ❓ What is a DAG? ❓ What are Jobs, Stages & Tasks? ❓ What is Lazy Evaluation? ❓ How do Executors p…
Asheesh T.
Trained 300+ Data Engineers | New Batch Starting Soon | DM to know more |
Must read 💯 A great resource for understanding ETL concepts, modern data pipelines, data quality and scalable data architectures. If you're looking to strong your fundamentals in data engineering I will recommend r…
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
𝗔𝘇𝘂𝗿𝗲 𝗶𝘀 𝘂𝘀𝗲𝗱 𝗯𝘆 𝟵𝟱% 𝗼𝗳 𝗙𝗼𝗿𝘁𝘂𝗻𝗲 𝟱𝟬𝟬 𝗰𝗼𝗺𝗽𝗮𝗻𝗶𝗲𝘀. 𝗧𝗵𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 𝘄𝗵𝗼 𝘄𝗼𝗿𝗸 𝘄𝗶𝘁𝗵 𝗶𝘁 𝗲𝗮𝗿𝗻 𝘀𝗼𝗺𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗵𝗶𝗴𝗵𝗲𝘀𝘁 𝘀𝗮𝗹𝗮𝗿𝗶𝗲𝘀 𝗶𝗻 𝘁𝗲𝗰𝗵. He…