Apache Spark
69%21 posts analysés sur les 12 dernières semaines
Sur les 12 dernières semaines
Moyenne tous posts confondus
Ce mois-ci vs le précédent
21 posts analysés sur les 12 dernières semaines
Sur les 12 dernières semaines
Moyenne tous posts confondus
Ce mois-ci vs le précédent
Pente de progression (6 semaines)
Meilleure semaine : 20 juil. (560 likes moy.)
Matei Zaharia
CTO & Cofounder at Databricks, CS Professor at Berkeley
Lots of new goodies in Apache Spark 4.2: Metric Views, GEOGRAPHY and GEOMETRY types, AutoCDC in Spark Declarative Pipelines, and more, from 1900 total commits. Available today in Databricks: https://lnkd.in/gmiaA73G
Ajay Kadiyala
Senior Data Engineer @ ENBD | Azure • Databricks • PySpark • SQL | I simplify Data Engineering, interviews & real-world project learning for 200K+ professionals
“Should I learn Scala or PySpark for Data Engineering?” I get this question repeatedly. My answer is For most Data Engineers today, start with PySpark. But that does not mean Scala is outdated or that Scala is always …
Ajay Kadiyala
Senior Data Engineer @ ENBD | Azure • Databricks • PySpark • SQL | I simplify Data Engineering, interviews & real-world project learning for 200K+ professionals
𝗔𝘇𝘂𝗿𝗲 𝗶𝘀 𝘂𝘀𝗲𝗱 𝗯𝘆 𝟵𝟱% 𝗼𝗳 𝗙𝗼𝗿𝘁𝘂𝗻𝗲 𝟱𝟬𝟬 𝗰𝗼𝗺𝗽𝗮𝗻𝗶𝗲𝘀. 𝗧𝗵𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 𝘄𝗵𝗼 𝘄𝗼𝗿𝗸 𝘄𝗶𝘁𝗵 𝗶𝘁 𝗲𝗮𝗿𝗻 𝘀𝗼𝗺𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗵𝗶𝗴𝗵𝗲𝘀𝘁 𝘀𝗮𝗹𝗮𝗿𝗶𝗲𝘀 𝗶𝗻 𝘁𝗲𝗰𝗵. Her…
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
SQL vs PySpark: Stop choosing one. Learn how they work together. One of the biggest mistakes I see aspiring Data Engineers make is treating SQL and PySpark as competing skills. They're not. SQL teaches you how to thin…
Abhishek Agrawal
Senior Azure Data Engineer | Azure Databricks | Microsoft Fabric | Azure Data Factory | PySpark | Spark | Azure Synapse | SQL | Data Platform | Open to Germany 🇩🇪 & UAE 🇦🇪
🚀 The Ultimate PySpark Cheatsheet is Here! PySpark is one of the most essential skills for every Data Engineer—but remembering every function, optimization technique, and best practice isn't easy. That's why I created…
Serigne Faye
Data engineer | BI Engineer | #Google cloud innovators | Gcp certified
Spark ≠ PySpark… mais Spark + Python = PySpark 👉 Apache Spark : - Un moteur de calcul distribué open source. - Spécialement conçu pour le Big Data. - Ultra rapide et très flexible. - Écrit en Scala, mais prend en …
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
Most data engineers can write a df.write(), but very few can tell you exactly what happens under the hood . : : : If you want to move from writing basic queries to building production-grade, optimized data pipelines, yo…
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
Do you use PySpark... but still don't understand what happens after you click Run? If terms like these still confuse you: ❓ What is a DAG? ❓ What are Jobs, Stages & Tasks? ❓ What is Lazy Evaluation? ❓ How do Executors p…
Asheesh T.
Trained 300+ Data Engineers | New Batch Starting Soon | DM to know more |
Must read 💯 A great resource for understanding ETL concepts, modern data pipelines, data quality and scalable data architectures. If you're looking to strong your fundamentals in data engineering I will recommend r…
Shilpa Das
23k+Linkedin | Data Engineering Manager | Pepsico | Ex-Adobe | Data Career Mentor| Databricks(2x) | DM for Collab |Youtube |☎️topmate.io/shilpa_das10
𝗔𝘇𝘂𝗿𝗲 𝗶𝘀 𝘂𝘀𝗲𝗱 𝗯𝘆 𝟵𝟱% 𝗼𝗳 𝗙𝗼𝗿𝘁𝘂𝗻𝗲 𝟱𝟬𝟬 𝗰𝗼𝗺𝗽𝗮𝗻𝗶𝗲𝘀. 𝗧𝗵𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 𝘄𝗵𝗼 𝘄𝗼𝗿𝗸 𝘄𝗶𝘁𝗵 𝗶𝘁 𝗲𝗮𝗿𝗻 𝘀𝗼𝗺𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗵𝗶𝗴𝗵𝗲𝘀𝘁 𝘀𝗮𝗹𝗮𝗿𝗶𝗲𝘀 𝗶𝗻 𝘁𝗲𝗰𝗵. He…