🚫 Ongoing session enrollments are closed.
We are accepting only 100 enrollments for the August 2026 batch starting 1 Aug. Secure your spot before it fills up.
Python · SQL · PySpark — same problem each day, two approaches. Then build a complete DE project: Bronze → Silver → Gold, SCD Type 2, CDC, backfill & tests.
All 4 sessions as one complete bundle — 10% off included.
15 days of Python · SQL · PySpark problems, then the Build DE Project — each step builds on the last.
Master the Python topics tested in DE interviews — loops, strings, lists, dicts, sets, file I/O, APIs, DSA patterns and generators. 15 days of problems, one session.
For loops, enumerate, zip, HashMap, HashSet, Stack, Queue, Deque, heapq — used in real DE interviews
CSV, JSON, TXT handling; API pagination & JSON parsing; SQLite parameterized queries
Comprehensions, generators for large files, itertools.groupby, copy vs deepcopy
Every SQL topic tested in DE interviews — filters, strings, dates, aggregation, joins, window functions, CTEs and mixed problems. 15 days, one session.
Rank employees per region, top-N per group, running totals, previous-row comparisons
Manager lookup via self join, customers with no orders, all join types with real DE datasets
Multi-CTE pipelines, correlated subqueries, GROUPING SETS, ROLLUP, full mixed problems
Solve every SQL problem from the roadmap again — using PySpark DataFrame API and Spark SQL. Same 15-day problem set, distributed computing approach.
filter(), where(), groupBy().agg(), join(), Window — every SQL pattern mapped to PySpark
rank(), dense_rank(), lag(), sum().over(), ntile() — all 15 SQL window problems in PySpark
Broadcast joins, repartition vs coalesce, data skew, AQE, shuffle explained for interviews
Four weekend sessions — build a complete local DE project from scratch using Python, PySpark and SQL. Bronze ingestion → Silver cleaning → SCD Type 2 → Gold layer.
Read CSV files with PySpark StructType schema enforcement, write idempotent partitioned Parquet to data/bronze/
Business rule validation, deduplication, SHA-256 hash change detection, expire-and-insert SCD on dim_customers
Daily KPIs (orders, revenue, AOV), top products, idempotent backfill by date, interview walk-through prep
Kafka · Airflow · Databricks Delta Live Tables
Live mock rounds · Resume review · System design walkthroughs
"Nice explanation, easy to understand"
"You are doing great job, explaining every concept along with possible interview question & learning pace is good but small suggestion, your speaking pace is fast little bit."
"Overall, the topics covered in Python are good. My suggestion is to go a bit slower, because many of us are from non coding and non IT backgrounds. It would really help if each function and topic is explained with the flow, process, and real-world use cases like why we use it and where it’s applied. Overall, the feedback and teaching are good."
"This is not a way to teach and communication should happen both ways not one. Also as said will teach from basic then why when ask certain questions we got reply like if I teach from basic then it will take 6 month to complete python. Any chance that I can get refund? If yes please do refund."
"Course content was up to date and related to real buiness requirements."
"session was very helpful for understanding basic concepts"
"Session was very informative"
"Session 1 covered the project kickoff, architecture overview, and Azure setup. Good sense of what's ahead, but it felt more lecture-style than interactive, and a few doubts around the setup itself weren't fully cleared by the end."
"Explain services in detail."
"Good teaching , need little more depth in concept explaination"
"The session was Good, the explanation is fast please pause it in each step"
"Very nice project, easy to understand"
"Very nice content detailed content"
"Very nice content, detailed as well"
"The session was informative and well-structured"
"Very good teaching sql query solving are very much good"
"Aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
"you'll be able to download the notes and learning material for your selected session instantly."
"Session was very helpful and infoformative"
"It was a great Session and very helpful!"
"need to know about recording of sesions"
"* Teaching speaking pace is too fast. * Concepts are understand * Need more explanation on why functions and methods are used. * More practical Data Engineering examples are needed. * Focus should be on student understanding, not just covering topics. * Learning quality matters more than the number of hours."
"Good note and explanation"
"Very good session and practice questions."
"Very good session and practice quetions"
"Very Good Session and good practice quetions."
"You have good knowledge and understanding of the subject. To make an even stronger impact, focus on improving your English communication and public speaking skills. Try to engage more actively in discussions, ask questions when needed, and avoid repeating the same words frequently. This will help us understand the concept more clearly."
"The way of explaining the concept is areally good and in simple way and the way to solve and approach a problem will tell us how we can use the concept thanks hariom to develop this type of content it's really good and helpful"
"amazing session which can help in real time projects"
"great resources and Explanation quality is good"
"Very nice explanation, every topic is made easy and taught ."
"Session is good , but agr hindi me hota to aur accha hota"
"Great content, clear explanations, and excellent patience in resolving doubts. The practical examples along with theory make learning easier. More focus on real-world problem-solving and interview-based scenarios will help students become job and interview ready. Overall, the Python, SQL, and PySpark sessions are very useful and well-structured"
"Good session. Never attended this type of sessions in my career. Trainer is helping very much to understand the concept easily"
"Nice material and expalination"
"It’s the best thing I joined so far sir is really good and have explained everything so nicely will recommend him to all other peers of mine. This one course will help me a lot to upskill"
"The class was amazing. I learned a lot, and it was really helpful for data engineer interviews."
"Excellent content with examples"
"The content is excellent with adequate depth and examples required for mastering Python for DE roles."
"As yeueuehebsiwijwywygwgwbbajisjanbshddhshsh"
"class was good and its really helpful and useful"
"Good interview preparation material"
"Good and informative session"
"The topics covered were excellent and well-structured. Communication was clear and effective throughout the sessions. One-to-one discussions were very helpful in resolving doubts and improving understanding. The teaching style is engaging and practical. Unlike many trainers who rely on a support team, you personally conduct the sessions, which adds significant value to the learning experience."
"Really nice sessions, thanks for highlighting the important concepts and interview topics."
"Great session, got to know many new concepts"
"great explanation and quality learning material"
"It was good, should have included tough examples and questions"
"Good session on SQL which was explained in a good manner"
"Good learning material. Explain very well."
"Great session. The notes were exceptionally well-prepared and provided an in-depth look at every topic."
"Good learning material. Explain very well."
"Very good content and explanation"
"It's a very detailed course with excellent coverage of every topic. The content is well-structured and easy to follow. Overall, it's a great course and definitely worth the investment"
"Is was a good session and I understand everything"
"It's a very detailed course with excellent coverage of every topic. The content is well-structured and easy to follow. Overall, it's a great course and definitely worth the investment"
"Learned about sql optimization, sub queries and window functions helped me to understand better with the leet code problems"
"It was good but keep entire video in one it is confusing"
"Session was very helpful for me to learn python. I really believe this will help me to improve my python skills"
"One of the best to crack the Python interview. Need to do separate session for python which should be zero to hero kind of thing"
"One of the best to crack SQL interview"
"Love the way that you are teaching"
"Love the way you are teaching"
"It is avery good session , a bit faster but when we go through learning material it will be good"
"It is a good sesssion and have most topics are covered"
"Amazing learning, good explaination"
"It was good session with practical"
"It was a great Session and very helpful!"
"Questions are good , keep the list updated"
"Session is good but writing query and explaining will be more efficient"
"The SQL training session was good, and the concepts were explained clearly. The instructor demonstrated good knowledge of the subject and covered the topics effectively. One suggestion for improvement would be to organize the study materials and notebooks in a proper sequence. Currently, some of the materials appear to be out of order, making it slightly difficult to follow along and refer back to the content during practice. A structured arrangement based on the training flow would enhance the learning experience. It would also be helpful if the required software, tools, and installation prerequisites were communicated in advance before the sessions. This would allow participants to complete the setup beforehand and avoid spending class time on installations or troubleshooting. Including more hands-on exercises and practical SQL scenarios would further strengthen the understanding of concepts and their real-world applications. Overall, the training was informative, the instructor explained the topics well, and the content was useful. With better organization of study materials and advance communication of software requirements, the learning experience would be even more effective."
"Not yet watched need time to watch soon I'll thanks"
"Need more important on explanation, it should be two way communication. You are explaining something then it should be clear to all students you are just wanted to complete whole syallbus.. You have used claude to generate docs and all notebook thats fine but work on your communication or else you can choose language which you are good and communicate well and address things well. No negative things from my end just a review just take it as positively"
"Didn't attend. Will go through recording."
"It's was good very useful topics covered. Highly recommend for others who wants to join."
"There is no proper communication. in the link it’s mentioned at 6pm to 12am but the class started at 2pm. There is no recordings provided yet to check"
"Good learning from scratch"
"All good everything is good"
"Good and engaging session"
"The session was very useful engaging and understandable"
"Good learning experience with clear and useful content."
"The session was great refresher"
"Good lecture and great explanation of topics"
"The session was awesome, had a good understanding of what questions are important to prepare for data engineering role interviews."
"Great session and very interactive"
"Very good and well deatails session"
"Ita very nice and according to the Data enginner perspective."
"The Session was good, it was helpful"
"The content and explanation was great."
"Gguijyggewrghhjjiiuytrewerggghjhhvcfddwwrtt"
"content is good nd informative."
"It was really nice and very helpful to begin with. Good starting point for the beginners."
"Demonstrated understanding of Python syntax and core concepts. Used appropriate data structures (lists, dictionaries, sets, etc.). Wrote readable and maintainable code."
"Good session with practical knowledge."
"Thank you for conducting the training session. The content was well-structured and easy to understand. The trainer explained concepts clearly, answered questions effectively, and kept the session engaging. The practical examples were particularly helpful."
"Great session amazing teaching"
"This course was informative. Thanks Team."
"Good structured notebook. Session was not that well organised But the content was good"
"All sessions are good,nice meterial and nice explanation as well"
"Good. I like the course content covered."
"Good Content for beginner's"
"The live session was excellent and highly insightful. The concepts were explained clearly with practical examples, making complex topics easy to understand."
"Worth it But divide sesions in 2 hours 3 session"
"Loved how structured python course was, and the Jupyter Notebooks with all the examples for every data engineering scenario were awesome 🤩. Keep helping us with great content brother🙌."
"Try to explain the things you've mentioned in details , not just upr upr se Sometime apne srf dikha diya ki notes me ye likha hai. Just a genuine feedback"
"Good and useful for DE interviews to understand fundamentals"
"very nice and interactive lecture"
"I had a wonderful experience with in detail explanation of every important topic."
"It was good and needed to share resources before even session start"
"Good explaination and very helpful"
"Python fundamentals data structures, pandas for cleaning, reshaping, merging large datasets, api calls"
"Thank you for the great session. I found the content incredibly productive and valuable. To help optimize future sessions, I wanted to share a couple of observations. While the high productivity was excellent, incorporating a few brief breaks would be highly beneficial to help maintain that level of focus and energy throughout. Additionally, it felt like a few complex topics were rushed through. Spending a bit more time on those areas and providing deeper context would make the information much easier to absorb. Thank you again for your time and effort!"
"Session was good but ,it is speedy. But overall it is good"
"great resources and good explanation"
"Very good and informative sessions"
"Content is good but explanation can be improved."
"Hellow Sir really appreciate your efforts… Speaking continuously for hours for us and really good detailed explanation… and making sure to explain everything… providing recordings notes long session in such reasonable price 🙏🏼🙏🏼🙏🏼🙏🏼"
"Its was good. If it better to divide in part for more connectivity."
"The course is very good and the notes were perfectly explained."
"Nice and helpful to learn"
"it has been a great learning experience so far. The concepts are explained in a clear and structured manner with practical examples, making them easy to understand. The session is interactive, and the trainer addresses questions patiently. I believe this course is very helpful for anyone looking to strengthen their Python skills for a Data Engineering career."
"It was good but i think two way communication is necessary"
"very nice and interactive session"
"Very very good learning session good teaching"
"It was good, but it was very fast need to spend some time on each topics"
"Good notes . Scenarios were good. Worth my money"
"1. No doubt it was good, but it had too many things which were not even useful, it is good to include them as a complete package. Put comments like not useful, just include here for practice... 2. Examples were too basic, include tough examples, real life interview ques, application based ques 3. Explaining speed was fast 4. Rather doing it for 6 hours, could do it for 3 hrs 3 days, it would be much better"
"The class is quote fast paced but notes are very helpful . Take one topic and explain slowly so we can also remember the concepts now we have to spend time to read each concept again"
"Session is good but after explaining concept give us pratice questions it will help us understand deep after explaining write for code on notebook for few secenerio so we can understand how you are approaching the problem by applying concept"
"Good session and informative want recording did not attending the session and want to learn fast about ghe session"
"It was a valuable session"
"Overall, the training content and study material are good, and the instructor explains the concepts well, which helps keep the sessions engaging rather than monotonous. One area for improvement would be to allocate more time for practical demonstrations and hands-on exercises alongside the theoretical concepts. Practical implementation helps reinforce learning and improves understanding of real-world applications. It would also be beneficial to address participant questions and clear doubts after completing each topic before moving on to the next. While I understand the time constraints, better planning of the sessions could help maintain a medium-paced learning approach rather than covering topics too quickly. Additionally, including more scenario-based discussions, real-world use cases, and interview-oriented questions would add significant value. As someone with 13 years of experience transitioning into Data Engineering, I would appreciate more guidance on industry expectations, the skills organizations look for in experienced professionals, and the types of technical and scenario-based interview questions commonly asked for Data Engineering roles. Overall, the teaching methodology is good, the content is relevant, and the sessions are informative. With additional focus on hands-on learning, doubt clarification, practical scenarios, and career guidance, the learning experience would be even more effective."
"the format and the way material is structured is helpful.it would be better a GOOGLE DRIVE link that contains the python materials for interview is shared inspite of sharing the recording in whatsapp for future use."
"It was good session but I had lost my knowledge about python completely so it was really complex for me . Though with practice I’ll learn more Session should be at least 2 so that we can cover & learn appropriately"
"I really enjoyed this Python tutorial. The concepts were explained in a simple and easy-to-understand way, and the practical examples made learning much easier. The hands-on exercises helped me apply what I learned and build confidence in coding. Overall, it was a great learning experience and has motivated me to continue improving my Python skills."
"most suggested to join and it's worth it in depths concept's concept clearing"
"Very Good Study Material and Good Pace of Teaching"
"Content was good and explained in a very good way"
"It is really good"
Everything you need to know about the Live Sessions
Loop over products with enumerate, pair names with salaries via zip, running total with break
Parse CSV rows, extract email domains, parse raw log lines into structured dicts
Second highest without full sort, deduplicate preserving order, daily temperatures stack problem
Two-sum with HashMap, max frequency word, group employees by department
Set difference on product IDs, first non-repeating character, intersect refunds and purchases
Mutation bug with =, fix with copy(); deepcopy nested dicts; override test pipeline config safely
csv module aggregation, filter JSON by status, generator for large log files line by line
Create table, insert rows, parameterized SELECT, bulk-insert with executemany + INSERT OR IGNORE
Flatten paginated API responses, generator pagination, filter valid records from mixed API results
Unique pairs summing to target, max 7-day revenue window, busiest 1-hour timestamp window
Balanced brackets, sliding window max with deque, event count in last 60 seconds
Top-5 products with heapq.nlargest, top-10 severity log events, median salary per department
Read CSV in 1000-row chunks, merge duplicate keys via dict comprehension, groupby totals
Invert a dict, customers with total spend > 500, flatten deeply nested order-item JSON
defaultdict grouping, copy vs deepcopy bug fix, generator for 1M-row CSV with status filter
Salary range filters, dept averages with HAVING, name-starts-with OR hire-year conditions
Standardize emails, split full_name into first/last, clean phone numbers with REPLACE
Delivery days (ship_date - order_date), monthly revenue groups, customers with first order in 90 days
Revenue per region, products with 50+ units sold, min/max/avg/total per region
Revenue by region+category, ROLLUP grand totals, GROUPING SETS multi-dimension in one query
Orders with customer names, customers with no orders (LEFT + IS NULL), totals including zero-order customers
Employee-manager via self join, FULL OUTER for unmatched products and orders, CROSS JOIN combinations
Deduplicated regional tables, year-over-year revenue comparison, 2024-not-2023 customers via EXCEPT
Rank employees by revenue per region, top-2 with tie handling, find rank-#1 employees
Previous order amount via LAG, running customer total, flag 2× spike orders
Customers above average via CTE, most recent order via correlated subquery, top-10 revenue with multi-CTE
3rd highest salary per dept, managers with large teams above avg salary, min/max product per region
Total spend + rank + percentile, NTILE top-25% flag, monthly cohort cumulative new customers
Salary vs dept avg + percentile rank, Active/At Risk/Churned classification, monthly revenue trend
ROW_NUMBER top earner per dept, EXCEPT for 2024-not-2023, 30-day rolling revenue with window functions
Same as SQL Day 1: salary range, HAVING equivalent, name/date conditions in PySpark
Same as SQL Day 2: standardize emails, split names, clean phone numbers in DataFrame API
Same as SQL Day 3: delivery days, monthly revenue groups, 90-day first-order filter
Same as SQL Day 4: revenue per region, top products, min/max/avg per region
Same as SQL Day 5: GROUPING SETS, ROLLUP, multi-level aggregation in PySpark
Same as SQL Day 6: orders with customer names, null-join for no-order customers
Same as SQL Day 7: manager lookup, CROSS JOIN combinations, unmatched records
Same as SQL Day 8: UNION ALL, year-over-year, EXCEPT equivalent via subtract()
Same as SQL Day 9: rank per region, top-2 with ties, rank-#1 employees
Same as SQL Day 10: previous order via lag(), running total, 2× spike flag
Same as SQL Day 11: CTEs via Spark SQL or intermediate aliases, correlated subquery equivalent
Same as SQL Day 12: 3rd highest salary, manager team filters, min/max product per region
Same as SQL Day 13: percentile via ntile, cumulative window, cohort report
Same as SQL Day 14: datediff(), when().otherwise() for classification, lower().contains() for name filter
Write Window.partitionBy().orderBy() from scratch; explain broadcast join, data skew mitigation
4 PM – 7 PM IST · Source CSV files with dirty data; ingest.py reads with StructType; write partitioned Parquet to data/bronze/; idempotent by date partition; log row counts in/out
Re-running same date overwrites only that partition; StructType catches bad data at the boundary; no silent nulls
4 PM – 7 PM IST · clean.py: drop NULL customer_id, cast order_date to DateType, standardize status (lowercase+strip), deduplicate by order_id keeping latest
amount <= 0 → is_valid=False, keep row for audit trail — do NOT drop; output DQ report: rows in/dropped/flagged/nulls per column
4 PM – 7 PM IST · scd_customers.py: SHA-256 hash of email/phone/address for change detection; expire old row (end_date=today, is_current=False); insert new version; print new/changed/unchanged/expired counts
Same SCD pattern applied to products; add price_change_flag when unit_price changes — same production pattern, different dimension
4 PM – 7 PM IST · daily_kpis.py: per-day total orders, total revenue, unique customers, average order value; write to data/gold/ partitioned by report_date
product_summary.py: top 10 products by revenue, orders per product — the aggregation pattern used in every DE take-home assessment
--start-date / --end-date CLI flags to reprocess specific date partitions; test idempotency by running same date twice — output must be identical
Verbal template: Bronze layer → why idempotent → Silver rule → SCD design decision → Gold metric → one challenge faced → what you would change at 100× scale