[Spring 2024] Beyond SQL: Dataframes in the Database (Devin Petersohn)
Dataframes are popular tools for interacting with and exploring data, but they are not as well understood nor as deeply studied as databases. Python's pandas. and Apache Spark are two of the most popular dataframes in use by data practitioners, but even these are extremely different from each other in terms of guarantees and user expectations. In this talk, we will explore these differences and take a deep dive into pandas-like dataframes with a theoretical lens, exploring the dataframe data... [Read More]