Generated by All in One SEO v4.9.10, this is an llms.txt file, used by LLMs to index the site. # Jesse Anderson ## Sitemaps - [XML Sitemap](https://www.jesse-anderson.com/sitemap.xml): Contains all public & indexable URLs for this website. ## Posts - [Gemini Batch API for Java](https://www.jesse-anderson.com/2025/11/gemini-batch-api-for-java/) - There isn't any documentation available about the Batch API for Gemini in Java. There are a few code samples, but not much explanation. Additionally, some of the API's advanced features lack examples as well. Due to the lack of documentation, it took significantly longer to create this example. In this post, I'll share the code - [Unapologetically Technical Episode 20 - Shane Murray](https://www.jesse-anderson.com/2025/05/unapologetically-technical-episode-20-shane-murray/) - Blog Summary: (AI Summaries by Summarizes)Shane Murray transitioned from studying math and finance in Sydney to leading AI strategy at Monte Carlo Data in New York.He pioneered online multivariate experimentation long before A/B testing became mainstream, with notable examples from various industries.Shane discusses the digital transformation at The New York Times, emphasizing the shift from - [Unapologetically Technical Episode 18 - Adrian Woodhead](https://www.jesse-anderson.com/2025/03/unapologetically-technical-episode-18-adrian-woodhead/) - Blog Summary: (AI Summaries by Summarizes)Adrian Woodhead is a distinguished software engineer and a pioneer in the European Hadoop ecosystem.He authored a chapter in 'Hadoop: The Definitive Guide', showcasing his expertise in data engineering.Adrian's career journey spans from South Africa to the Netherlands, influenced by the dot-com boom.He was involved in pioneering work with touch - [Unapologetically Technical Episode 16 - David Jayatillake](https://www.jesse-anderson.com/2025/01/unapologetically-technical-episode-16-david-jayatillake/) - Blog Summary: (AI Summaries by Summarizes)David Jayatilake is the Head of AI at Cube, with extensive experience in analytics and AI solutions.He progressed from an analyst to a pricing manager at Worldpay, optimizing processes through data engineering.David emphasizes the importance of aligning data teams with business needs to drive value.He has leadership experience at Elevate - [Unapologetically Technical Episode 17 - Semih Salihoglu](https://www.jesse-anderson.com/2025/02/unapologetically-technical-episode-17-semih-salihoglu/) - Blog Summary: (AI Summaries by Summarizes)Semih Salihoglu is an Associate Professor at the University of Waterloo and co-founder of Kuzu, an embedded graph database.He has a background in distributed systems and databases, with significant industry experience from Google.Salihoglu's academic journey includes studying computer science and economics at Yale University and pursuing doctoral studies at Stanford - [Unapologetically Technical Episode 19 - Jacopo Tagliabue](https://www.jesse-anderson.com/2025/04/unapologetically-technical-episode-19-jacopo-tagliabue/) - Blog Summary: (AI Summaries by Summarizes)Jacopo Tagliabue, founder of Bauplan, shares insights from his journey in data science and entrepreneurship.He emphasizes the evolution of data work from rule-based NLP to the transformative power of LLMs.Bauplan's architecture focuses on innovative functions as compute units and efficient data versioning using Git.Jacopo discusses the importance of fast feedback - [Unapologetically Technical Episode 15 - Frances Perry](https://www.jesse-anderson.com/2024/12/unapologetically-technical-episode-15-frances-perry/) - Blog Summary: (AI Summaries by Summarizes)Frances Perry, Head of Engineering at Mother Duck, shares insights from her 16 years at Google and discusses challenges and rewards of scaling engineering teams.The conversation explores the future of data processing with DuckDB and Mother Duck, emphasizing the potential of single-node databases and more efficient data solutions.Frances Perry highlights - [Unapologetically Technical Episode 14 - Cliff Crosland](https://www.jesse-anderson.com/2024/10/unapologetically-technical-episode-14-cliff-crosland/) - Blog Summary: (AI Summaries by Summarizes)Cliff Crosland, CEO of Scanner.dev, emphasizes the importance of effective log analysis in today's complex systems, viewing logs as a treasure trove of insights.Crosland's expertise in distributed systems, graph creation, and entity resolution provides valuable insights into the implications of Generative AI and LLMs for current and future coders.The challenges - [Data Teams Survey 2020-2024 Analysis](https://www.jesse-anderson.com/2024/09/data-teams-survey-2020-2024-analysis/) - Blog Summary: (AI Summaries by Summarizes)**Total Value Creation**:**Gradual Decrease in Value Creation**:**Team Makeup and Descriptions**:**Methodologies**:**Advice**: Survey Changes Over Time Between 2020 and 2024 (see 2020, 2023, and 2024 for each year’s information), I’ve been conducting a data teams survey. I wanted to dedicate an entire post to examining the change in data teams over time. - [Data Teams Survey 2024 Results](https://www.jesse-anderson.com/2024/08/data-teams-survey-2024-results/) - Blog Summary: (AI Summaries by Summarizes)Companies are not fully utilizing LLMs in data engineering, with 24.7% of teams not using them at all.Only 12% of teams are using LLMs for data processing, the most ideal use.Challenges in using LLMs include concerns about human-generated data, costs, and long response times.Data science, engineering, and operations are essential - [Unapologetically Technical Episode 13 - Jeff Chou](https://www.jesse-anderson.com/2024/08/unapologetically-technical-episode-13-jeff-chou/) - Blog Summary: (AI Summaries by Summarizes)Jeff Chou, CEO and co-founder of Sync Computing, shares his journey from academia to startup life, highlighting how his experience with simulations shaped the vision for Sync Computing.The challenges of translating academic research into real-world solutions are explored in this episode.Bridging the gap between technical execution and business objectives is - [Unpacking the Latest Streaming Announcements: A Comprehensive Analysis](https://www.jesse-anderson.com/2024/06/unpacking-the-latest-streaming-announcements-a-comprehensive-analysis/) - Blog Summary: (AI Summaries by Summarizes)**Using Kafka Protocol**: The value of Kafka lies in its protocol, which will outlive Apache Kafka.**Competition in the Space**: Increasing competition expected from other vendors, with support for the Kafka protocol being a key factor.**Cost Trade-offs**: When creating streaming systems, choosing between cost, availability, and performance is crucial.**Multi-tenancy and Cost**: - [Unapologetically Technical Episode 12 - AJ Hunyady](https://www.jesse-anderson.com/2024/06/unapologetically-technical-episode-12-aj-hunyady/) - Blog Summary: (AI Summaries by Summarizes)Interview with AJ Hunyady, founder and CEO of InfinyOn.Discussion on early experiences with networking systems and how they prepared him for data work.Implications of Generative AI and LLMs for current and future coders.Challenges of using batch systems in security and the need for real-time actions.Views on containerization and Kubernetes consolidation - [Unapologetically Technical Episode 11 - Hubert Dulay](https://www.jesse-anderson.com/2024/05/unapologetically-technical-episode-11-hubert-dulay/) - Blog Summary: (AI Summaries by Summarizes)Hubert Dulay, author of "Streaming Data Mesh" and Developer Advocate at StarTree, shares insights on transitioning from web development to data work.Dulay discusses his experiences with web backends like CORBA and SOAP and how they prepared him for working with data.The interview covers Dulay's time at Cloudera and Confluent, highlighting - [Unapologetically Technical Episode 10 - Michael Drogalis](https://www.jesse-anderson.com/2024/04/unapologetically-technical-episode-10-michael-drogalis/) - Blog Summary: (AI Summaries by Summarizes)Interview with Michael Drogalis, founder and CEO of Shadow Traffic, discussing the early Hadoop era and the importance of Kafka in the industry.Insights shared on starting a new company in his 20s and being acquired by Confluent.Discussion on how software engineers can enhance their projects by understanding product creation and - [The Difference Between Learning and Doing](https://www.jesse-anderson.com/2023/11/the-difference-between-learning-and-doing/) - Blog Summary: (AI Summaries by Summarizes) Learning options trading involves data and programming but is not as technical as data engineering or software engineering. Different types of learning videos include Hype, Low Effort, Novice, and Professional categories. Watching hype, low-effort, and novice videos in options trading can be detrimental as they often lack substance and - [Why Most Data Projects Fail & How to Avoid It at GOTO 2023](https://www.jesse-anderson.com/2024/03/why-most-data-projects-fail-how-to-avoid-it-at-goto-2023/) - Blog Summary: (AI Summaries by Summarizes)Data projects often fail due to similar reasons, with many management and data teams lacking understanding of project success factors.Understanding the "who, what, when, where, and how" of data projects can significantly improve team outcomes.Addressing key questions is crucial to avoid project failures.The speaker presented on "Why Most Data Projects - [Unapologetically Technical Episode 9 - Gunnar Morling](https://www.jesse-anderson.com/2024/02/unapologetically-technical-episode-9-gunnar-morling/) - Blog Summary: (AI Summaries by Summarizes)Gunnar Morling, creator of the Billion Row Challenge and Senior Staff Software Engineer at Debezium, emphasizes the importance of staying in a position long enough to gain experience and witness the success or failure of decisions.Gunnar Morling shares his experiences at Red Hat and working on Debezium during the interview.The - [The State of Data Engineering at Data Day Texas 2024](https://www.jesse-anderson.com/2024/01/the-state-of-data-engineering-at-data-day-texas-2024/) - Blog Summary: (AI Summaries by Summarizes)The premier of the latest talk covering The State of Data Engineering, discussing the industry's history and future direction.The talk starts with data warehousing and progresses into data science, highlighting key insights and trends.The importance of data engineering in avoiding the same fate as data warehousing and data science is - [Unapologetically Technical Episode 8 - Tom Scott](https://www.jesse-anderson.com/2024/02/unapologetically-technical-episode-8-tom-scott/) - Blog Summary: (AI Summaries by Summarizes)Tom Scott, Founder and CEO of Streambased, is interviewed in the latest episode of Unapologetically Technical.The discussion covers distributed systems and the creation of Monte Carlo simulations.Tom Scott's work history includes creating a data warehouse at Sky Betting, working at Cloudera in Customer Operations Engineering, and contributing to projects at - [Unapologetically Technical Episode 7 - Stephane Derosiaux](https://www.jesse-anderson.com/2023/12/unapologetically-technical-episode-7-stephane-derosiaux/) - Blog Summary: (AI Summaries by Summarizes)New episode of Unapologetically Technical featuring an interview with Stephane Derosiaux from ConduktorDiscussion on evolving architectures and creating real-time systems at Auchan (grocery) and Adeo/Leroy Merlin (Home Improvement)Issues of British food and tips on finding good food in LondonConversation about geeking and views on minimalismDeep dive into Conduktor and how - [Unapologetically Technical Episode 6 - Matteo Merli](https://www.jesse-anderson.com/2023/11/unapologetically-technical-episode-6-matteo-merli/) - Blog Summary: (AI Summaries by Summarizes)Matteo Merli, co-creator of Apache Pulsar and CTO of StreamNative, discusses the creation of Apache Pulsar and its significance.The reasons behind creating Apache Pulsar at Yahoo and how new projects were managed and convinced.The differences between US and European technology companies and Matteo's experiences with the transition.Detailed insights into Apache - [Unapologetically Technical Episode 5 - Neil Avery](https://www.jesse-anderson.com/2023/10/unapologetically-technical-episode-5-neil-avery/) - Blog Summary: (AI Summaries by Summarizes)Neil Avery from Liquidlabs shared insights on creating grid computing systems at major banks and founding a startup called Logscape.The discussion also covered the early days of Confluent, including topics like ksqlDB, NIH, and what makes a successful startup in its early stages.SpecMesh, a tool discussed in detail, was highlighted - [Unapologetically Technical Episode 4 - Karthik Ramasamy](https://www.jesse-anderson.com/2023/07/unapologetically-technical-episode-4-karthik-ramasamy/) - Blog Summary: (AI Summaries by Summarizes)Karthik Ramasamy, head of Streaming at Databricks, discusses Spark Structured Streaming in-depth in the latest episode of Unapologetically Technical.The episode covers Karthik's experience at the University of Wisconsin-Madison and the impact it had on creating distributed systems thinkers.Karthik also shares insights from his time at Twitter, Streamlio, and Splunk.Unapologetically Technical - [Unapologetically Technical Episode 3 - Abhishek Chauhan](https://www.jesse-anderson.com/2023/06/unapologetically-technical-episode-3-abhishek-chauhan/) - Blog Summary: (AI Summaries by Summarizes)Abhishek Chauhan, CTO and Co-founder at Grainite, discusses early adoption of new technologies and his experiences at Sun Microsystems and Citrix.Chauhan's experience with transactions and reliability at Citrix led him to create Grainite, focusing on a better way of handling these aspects.Detailed discussion on the current exactly-once systems and their - [Unapologetically Technical Episode 2 - Tomas Neubauer](https://www.jesse-anderson.com/2023/05/unapologetically-technical-episode-2-tomas-neubauer/) - Blog Summary: (AI Summaries by Summarizes)Episode 2 of Unapologetically Technical features Tomas Neubauer, CTO and Co-founder at Quix, discussing his time at McLaren and how McLaren creates data products for F1 racing.Tomas Neubauer shares the issues and best practices he used while creating data products at McLaren.Quix makes it easier to create real-time processing and - [Unapologetically Technical Episode 1 - Roy Hasson](https://www.jesse-anderson.com/2023/04/unapologetically-technical-episode-1-roy-hasson/) - Blog Summary: (AI Summaries by Summarizes)Roy Hasson, Senior Director of Product Management at Upsolver, shares insights on creating better data products based on his experience at Amazon.Roy's background as a Sales Engineer at Amazon helped him develop better products, including running Product for Amazon Glue.The discussion covers the factors that contribute to the success or - [The Data Discovery Team](https://www.jesse-anderson.com/2023/11/the-data-discovery-team/) - Blog Summary: (AI Summaries by Summarizes)Data discovery team plays a crucial role in searching for data in the IT landscape.Data discovery team must make data discoverable to the operations team.Collaboration between data discovery team and operations team is essential for creating data products.Strong domain knowledge and technical skills are required for effective data discovery team - [Current 2023 Announcements](https://www.jesse-anderson.com/2023/10/current-2023-announcements/) - Blog Summary: (AI Summaries by Summarizes) Confluent announced the addition of serverless Flink, expanding their moats to three, alongside replication and Confluent Cloud. Confluent’s marketing tends to oversimplify complex topics like the universal use of Kafka and streaming. Kafka is not universally suitable, requiring a clear need for streaming over batch processing. While Confluent Cloud - [Data Teams Survey 2023 Follow-Up](https://www.jesse-anderson.com/2023/05/data-teams-survey-2023-follow-up/) - Blog Summary: (AI Summaries by Summarizes) Many companies are utilizing data mesh as a common methodology. Choosing a methodology requires understanding that there is no one-size-fits-all solution. Homegrown methodologies are prevalent, combining elements from various established methodologies. Implementing methodologies fully can be challenging, especially in data-related fields. Companies often struggle to fully implement agile methodologies, - [GPT and LLMs from a Data Engineering Perspective](https://www.jesse-anderson.com/2023/09/gpt-and-llms-from-a-data-engineering-perspective/) - Blog Summary: (AI Summaries by Summarizes) GPT and LLMs are gaining attention from data science and business perspectives, but less so from the data engineering side. Google is integrating LLMs into Google Workspace, making interaction with LLMs more seamless for users. The widespread adoption of LLMs will lead people to expect their presence in daily - [The Soldiers, Rogues, and Mages of Data Teams](https://www.jesse-anderson.com/2022/03/the-soldiers-rogues-and-mages-of-data-teams/) - Blog Summary: (AI Summaries by Summarizes) Data teams are like Role Playing Games (RPGs), where individuals work together for a common goal, each with their own levels, skills, and stats. Classes in RPGs equate to core abilities that characters are good at, such as soldiers, rogues, and mages, with each class complementing the others in - [Enabling The People, Enabling The Data with Kulani Likotsi](https://www.jesse-anderson.com/2022/11/enabling-the-people-enabling-the-data-with-kulani-likotsi/) - Blog Summary: (AI Summaries by Summarizes) Kulani Likotsi has had a successful career journey in data management and data governance, showcasing a growth mindset and a willingness to learn. Data privacy laws like GDPR and POPIA are crucial in data governance, emphasizing the importance of data protection and privacy. Data teams should initiate data governance - [Brief History of Data Engineering](https://www.jesse-anderson.com/2022/12/brief-history-of-data-engineering/) - Blog Summary: (AI Summaries by Summarizes) Google created MapReduce and GFS in 2004 for scalable systems. Apache Hadoop was created by Doug Cutting in 2005 based on Google’s papers. Cloudera and Hortonworks commercialized open-source big data technologies in 2008 and 2011. Apache Hive and Apache Pig were introduced to enhance Hadoop with SQL capabilities. Apache - [Analysis of Confluent Buying Immerok](https://www.jesse-anderson.com/2023/01/analysis-of-confluent-buying-immerok/) - Blog Summary: (AI Summaries by Summarizes) Customers should benefit as long as they don’t have significant investments in ksqlDB or perhaps Kafka Streams. The general purpose real-time compute is coalescing around Spark Streaming and Flink. Looking at this purchase, this isn’t a merger like Cloudera and Hortonworks. There are fewer choices now, but more mature - [Data Teams Survey 2023 Results](https://www.jesse-anderson.com/2023/03/data-teams-survey-2023-results/) - Blog Summary: (AI Summaries by Summarizes) Companies with all three teams (green) or data science and data engineering (medium blue) create the highest success. Best practices for ensuring data projects meet business needs include working with the business, continuous integration, having a qualified data engineering team, and leveraging automation. Respondents with the highest value creation - [Big Data and Analytics in the COVID-19 Era](https://www.jesse-anderson.com/2020/03/big-data-and-analytics-in-the-covid-19-era/) - Blog Summary: (AI Summaries by Summarizes)Data teams should focus on solving real business problems rather than just storing data.Creating models that optimize efficiency and save money is crucial, especially in the COVID-19 era.High-quality data is essential for successful model deployment.Organizations must ensure data security and proper handling of sensitive information, especially when working from home.Models - [Data Engineering Technology Tree](https://www.jesse-anderson.com/2020/12/data-engineering-technology-tree/) - Blog Summary: (AI Summaries by Summarizes) Data engineering requires in-depth knowledge of various technologies and big data. Visualizing skills as a technology tree can help in understanding the complexity of data engineering. Skipping foundational knowledge in data engineering can lead to problems similar to skipping technologies in a game like Civilization. Becoming a proficient data - [It's Time to Change How We Manage Data Teams](https://www.jesse-anderson.com/2021/02/its-time-to-change-how-we-manage-data-teams/) - Blog Summary: (AI Summaries by Summarizes) Spreading out a problem to more computers can leverage resources better and faster. Centralizing innovation and strategy on a small management team may limit diverse perspectives and solutions. Involving individual contributors more in planning can lead to better outcomes. Data teams replicating top-down innovation may overlook crucial insights from - [Why Data Science Teams Don't Think They Need Data Engineering](https://www.jesse-anderson.com/2021/04/why-data-science-teams-dont-think-they-need-data-engineering/) - Blog Summary: (AI Summaries by Summarizes) Data science teams may mistakenly believe they don’t need data engineering, leading to underperformance and technical debt. Lack of understanding about the role of data engineers can result in unrealistic expectations and underestimation of complexity. Creating repeatable data science processes is crucial for efficiency and maintenance of data products. - [Keeping Things Stupidly Simple With Pulsar and Kafka](https://www.jesse-anderson.com/2021/09/keeping-things-stupidly-simple-with-pulsar-and-kafka/) - Blog Summary: (AI Summaries by Summarizes) Transitioning to Pulsar from Kafka is recommended for companies looking to simplify their architecture and avoid potential data loss and latency issues. Designing with Pulsar can eliminate the need for complex workarounds like Uber’s Consumer Proxy and streamline processes. Understanding the technologies being used is crucial for deploying them - [Ten Years On - The Million Monkeys Project](https://www.jesse-anderson.com/2021/10/ten-years-on-the-million-monkeys-project/) - Blog Summary: (AI Summaries by Summarizes) Building personal projects requires clear goals and a high level of effort to achieve meaningful outcomes. Effort put into a project can impress the right people and showcase creativity and mastery. Mastery in technical aspects of a project can set you apart from competition during interviews. Overcoming self-doubt and - [Analysis of Confluent's S1](https://www.jesse-anderson.com/2021/06/analysis-of-confluents-s1/) - Blog Summary: (AI Summaries by Summarizes) Confluent’s S1 filing for IPO signifies a major milestone in their journey. Apache Kafka and Confluent have emerged as superior solutions for real-time streaming needs, with Kafka solving problems elegantly. Confluent’s products are centered around Apache Kafka, with additional proprietary software like ksqlDB, Control Center, REST Server, and Schema - [What Happens When Data Science Teams Add A Data Engineer](https://www.jesse-anderson.com/2021/06/what-happens-when-data-science-teams-add-a-data-engineer/) - Blog Summary: (AI Summaries by Summarizes) Organizations are recognizing the importance of data engineering, but there is often a misunderstanding that simply adding a data engineer to the team will solve all problems. Data science teams may not fully buy into the critical nature of data engineering, leading to challenges in collaboration and understanding the - [Friends Don't Let Friends Copy and Paste](https://www.jesse-anderson.com/2010/11/friends-dont-let-friends-copy-and-paste/) - I ran into this comment in some code and just had to laugh (and refactor it). - [KPIs Every Data Team Should Have](https://www.jesse-anderson.com/2022/03/kpis-every-data-team-should-have-3/) - Blog Summary: (AI Summaries by Summarizes)KPIs are crucial for data teams as they differ from other teams in terms of value creation and performance.Establishing baseline numbers for metrics is essential before diving into KPIs to track maturity and growth effectively.Breaking down KPIs to a project/product level may be necessary when analyzing projects or data products.Business-focused - [The Reasons for Data Mesh on Pulsar](https://www.jesse-anderson.com/2022/04/the-reasons-for-data-mesh-on-pulsar/) - Blog Summary: (AI Summaries by Summarizes)Data mesh is becoming a popular way for companies to implement their data strategy, involving both organizational and technical changes.Flexibility is a key aspect of data mesh, allowing for easier creation and consumption of data products with less coordination between teams.Apache Pulsar and Apache Kafka are compared for their publish/subscribe - [An In-Depth Data Mesh Discussion with Zhamak Dehghani](https://www.jesse-anderson.com/2022/06/an-in-depth-data-mesh-discussion-with-zhamak-dehghani/) - Blog Summary: (AI Summaries by Summarizes)Zhama’s concept of data mesh is a paradigm shift in managing data-driven value at scale.Data mesh surfaced as a solution to unspoken challenges, not just a hyped trend.The return on investment for implementing data mesh may take months or years to materialize.Complexity is a key factor in determining if data - [Provoking Consumer-First Analytical Thinking with Drew Smith](https://www.jesse-anderson.com/2022/07/provoking-consumer-first-analytical-thinking-with-drew-smith/) - Blog Summary: (AI Summaries by Summarizes)Drew Smith, Vice President of Global Data and Analytics at Little Caesars Enterprises and Ilitch Companies, has a diverse background working with companies like IKEA and the International Institute for Analytics.Drew ran the Analytics Leadership Consortium at the International Institute for Analytics, where analytics leaders discussed organizational models and challenges - [The Art and Science of Data Storytelling with Brent Dykes](https://www.jesse-anderson.com/2022/10/the-art-and-science-of-data-storytelling-with-brent-dykes/) - Blog Summary: (AI Summaries by Summarizes)Data storytelling is a structured approach for communicating insights using narrative elements and explanatory visuals.Effective data storytelling involves three key elements: data, narrative, and visuals.The narrative is crucial in organizing and structuring a data story, similar to a plot with rising action, climax, and resolution.Connecting findings and key points into - [See, Build, Test, Experiment: Using Data Science to Change the World with Erick Webbe](https://www.jesse-anderson.com/2022/11/see-build-test-experiment-using-data-science-to-change-the-world-with-erick-webbe/) - Blog Summary: (AI Summaries by Summarizes)Erick Webb, Head of Data Science at bol.com, emphasizes the importance of experimentation in data science and problem-solving.Webb suggests focusing on solving problems effectively rather than pursuing the most elaborate solutions.Building credibility by consistently solving problems is crucial for establishing a reputation in data science.Webb recommends starting problem-solving by adding - [Independent Anniversary](https://www.jesse-anderson.com/2022/10/independent-anniversary/) - Blog Summary: (AI Summaries by Summarizes)Founding Big Data Institute independently eight years ago marked a significant milestone in the journey towards establishing an independent big data consulting company.Independence allows for a unique perspective on technology and vendors, free from corporate biases and constraints.Lack of independent thought in technology discussions is prevalent, with content often skewed - [Streamline Data Gathering](https://www.jesse-anderson.com/2018/08/streamline-data-gathering/) - Lorem ipsum dolor sit amet, consectetur adipiscing elit. Senectus amet et erat at facilisi elementum. - [Data Teams Survey Results](https://www.jesse-anderson.com/2020/12/data-teams-survey-results/) - Blog Summary: (AI Summaries by Summarizes) Data teams are crucial for the success of data projects, requiring collaboration between data science, data engineering, and operations. Companies with all three teams or at least data science and data engineering teams show higher success rates in creating value from data projects. Friction, lack of individual contributors, and - [What It Looks Like When a Team Is Missing](https://www.jesse-anderson.com/2020/10/what-it-looks-like-when-a-team-is-missing/) - Blog Summary: (AI Summaries by Summarizes) Data teams require all of their parts to be complete and succeed. Organizations or team members often don’t understand the issues when a team is missing and may blame themselves or technology for perceived issues. A common misconception is that only one team (data science) is needed to do - [Announcement: Data Teams Is Out!](https://www.jesse-anderson.com/2020/09/announcement-data-teams-is-out/) - Blog Summary: (AI Summaries by Summarizes) **Data Teams Book**: A unified management model for successful data-focused teams is available for purchase, aiming to increase the percentage of successful big data projects. **Years of Work and Research**: The book represents extensive work and research, offering insights on starting and improving data teams based on real-world experience. - [Kafka's Got a Brand-New Poll](https://www.jesse-anderson.com/2020/09/kafkas-got-a-brand-new-poll/) - Blog Summary: (AI Summaries by Summarizes) Kafka 2.0 introduced a new `poll()` method that takes a `Duration` as an argument, replacing the previous `poll()` method that took a `long`. Overloaded methods should have similar functionality; for example, `System.out.println(char[])` and `System.out.println(String)` are convenience methods with the same output. The new `poll()` method in Kafka has different - [Saving Money on Data Engineering in the Cloud](https://www.jesse-anderson.com/2020/04/saving-money-on-data-engineering-cloud/) - Blog Summary: (AI Summaries by Summarizes) Proactively reducing cloud costs can make data engineering teams look good to the CFO and demonstrate cost-saving initiatives. Understanding the organization’s focus on speed versus cost optimization is crucial for data engineering teams to make informed decisions. Collaborating with business stakeholders to present cost-saving opportunities based on cluster optimization - [Why I Recommend My Clients NOT Use KSQL and Kafka Streams](https://www.jesse-anderson.com/2019/10/why-i-recommend-my-clients-not-use-ksql-and-kafka-streams/) - Blog Summary: (AI Summaries by Summarizes)Implementing checkpointing in processing frameworks like Flink is crucial for ensuring quick recovery and minimal downtime in distributed systems.The lack of checkpointing in Kafka Streams can result in hours of downtime in case of failures, impacting the reliability and availability of the system.Kafka Streams lacks proper checkpointing mechanisms, making it - [Getting Your Programming Skills Ready for Data Engineering](https://www.jesse-anderson.com/2019/06/getting-your-programming-skills-ready-for-data-engineering/) - Blog Summary: (AI Summaries by Summarizes) Preparation for the data engineering course began early, showcasing a proactive approach to learning and exploring programming interests. Learning programming through resources like “Introduction to Statistical Learning in R” and “Learn Python the Hard Way” provided valuable insights and foundational knowledge. Practical application of programming skills through projects like - [I Come Not To Bury Cloudera But To Praise It](https://www.jesse-anderson.com/2019/06/i-come-not-to-bury-cloudera-but-to-praise-it/) - Blog Summary: (AI Summaries by Summarizes) Cloudera and MapR are facing challenges impacting their market value, with Cloudera’s stock price closing at $5.21 on June 6, 2019. Major changes in the Hadoop distribution landscape are possible, potentially leading to acquisitions and restructuring. The value generated from big data projects often falls short of expectations, raising - [Reducing Operational Overhead with Pulsar Functions](https://www.jesse-anderson.com/2019/05/reducing-operational-overhead-with-pulsar-functions/) - Blog Summary: (AI Summaries by Summarizes) Moving from racking and stacking to virtual machines to containerization represents operational evolution. Operationalizing consumers, producers, and consumer/producers can be tedious and complex. Containers simplify process deployment and monitoring but still involve duplicated code. Containerization only partially solves the operational challenges, leaving room for improvement. Focusing on essential components - [Reducing System Complexity with Event Sourcing](https://www.jesse-anderson.com/2019/03/reducing-system-complexity-with-event-sourcing/) - Blog Summary: (AI Summaries by Summarizes) Adopting event sourcing patterns can significantly increase developer productivity by reducing system complexity. High system complexity can hinder value creation as developers spend more time on complexity issues. Loosely coupling systems can help reduce complexity and improve developer productivity. Paying down technical debt is crucial for addressing system complexity - [Saving Money with Apache Pulsar Tiered Storage](https://www.jesse-anderson.com/2019/02/saving-money-with-apache-pulsar-tiered-storage/) - Blog Summary: (AI Summaries by Summarizes) Companies can save up to 85% on overall storage costs with forward planning when rolling out real-time messaging systems. Understanding the differences in data storage between Apache Kafka and Apache Pulsar is crucial for cost comparisons. In Kafka, the Broker process handles all data movement and storage, while in - [Q and A: Viewpoints on Open Source](https://www.jesse-anderson.com/2019/01/q-and-a-viewpoints-on-open-source/) - Blog Summary: (AI Summaries by Summarizes) There are diverse viewpoints on open source and its usage as a service. There is a moral obligation to give back to open source if you’re directly commercializing the project. The reason a company or organization should give back to open source is to recruit, show they can support - [The Three Components of a Big Data Data Pipeline](https://www.jesse-anderson.com/2019/01/the-three-components-of-a-big-data-data-pipeline/) - Blog Summary: (AI Summaries by Summarizes) Big Data pipelines require components from three different general types of technologies: compute, storage, and messaging. Apache Spark is just one part of a larger Big Data ecosystem that’s necessary to create data pipelines. Batch data pipelines require solving two core problems: compute and storage of data. Real-time Big - [Advice for Small Teams and Startups on Data Engineering](https://www.jesse-anderson.com/2018/12/advice-for-small-teams-and-startups-on-data-engineering/) - Blog Summary: (AI Summaries by Summarizes) Small data engineering teams require different tactics than larger companies and teams. The first data engineering hire is crucial as they give the team its starting direction and make many of the initial technology decisions. Hiring competent people is crucial early on as a single wrong hire can make - [Amazon and Open Source Business Models and You](https://www.jesse-anderson.com/2018/12/amazon-and-open-source-business-models-and-you/) - Blog Summary: (AI Summaries by Summarizes) Open source companies like Cloudera and Confluent make money by selling support, training, consulting, and management tools for their free and open source projects. Amazon Web Services (AWS) makes money by creating managed services that make it easy for customers to use open source projects like Kafka and Hadoop. - [Creating a Data Engineering Culture](https://www.jesse-anderson.com/2018/11/creating-a-data-engineering-culture/) - Blog Summary: (AI Summaries by Summarizes) Data engineering culture is often implicit and assumed in organizations. Creating a data engineering culture involves recognizing the value and importance of data engineering at all levels of the organization. The right ratio of data scientists to data engineers is generally 2 to 5. Failure in big data projects - [Why You Can't Do All of Your Data Engineering with SQL](https://www.jesse-anderson.com/2018/10/why-you-cant-do-all-of-your-data-engineering-with-sql/) - Blog Summary: (AI Summaries by Summarizes) SQL cannot handle all aspects of creating a Big Data data pipeline, and eventually, a programming language will be needed to fill in the gaps. Customization is a key factor in the need for programming, as companies often have custom use cases that require custom code. Integrating systems is - [Thoughts on Cloudera Merging/Buying Hortonworks](https://www.jesse-anderson.com/2018/10/thoughts-on-cloudera-buying-hortonworks/) - Blog Summary: (AI Summaries by Summarizes) Cloudera has merged with/purchased Hortonworks, two fierce rivals in the data platform industry. There was a lot of animosity between the two companies, with an unspoken rule that you didn’t leave one for the other. The companies fought each other in the trenches of the Apache world, creating their - [Creating Work Queues with Apache Kafka and Apache Pulsar](https://www.jesse-anderson.com/2018/08/creating-work-queues-with-apache-kafka-and-apache-pulsar/) - Blog Summary: (AI Summaries by Summarizes) Kafka and Pulsar are commonly used for creating work queues, which involve publishing a message to be consumed by a cluster of processes for longer periods of processing time. Work queues require distributing processing across many different processes and computers, making the task more complex. Examples of work queues - [InfiniteConf Keynote - Why Real-time is the Future](https://www.jesse-anderson.com/2018/08/infiniteconf-keynote-why-real-time-is-the-future/) - Blog Summary: (AI Summaries by Summarizes) Real-time technology is gaining momentum in the business world. Real-time technology can provide businesses with valuable insights and help them make better decisions. Real-time technology can also benefit data sciences by allowing for faster and more accurate analysis of data. Common use cases for real-time technology include fraud detection, - [What is a Data Pipeline?](https://www.jesse-anderson.com/2018/08/what-is-a-data-pipeline/) - Blog Summary: (AI Summaries by Summarizes) Data pipeline is a collection of instructions to read, transform, or write data that is designed to be executed by a data processing engine. A data pipeline can be arbitrarily complex and can include various types of processes that manipulate data. ETL is just one type of data pipeline, - [Professional Data Engineering Review - Sanjoy Roy](https://www.jesse-anderson.com/2018/07/professional-data-engineering-review-sanjoy-roy/) - Blog Summary: (AI Summaries by Summarizes) The Professional Data Engineering course by Jesse Anderson covers a well-rounded curriculum that includes data engineering and data science. The course is a semester’s worth of learning and covers all relevant topics, including data ingestion, processing, and visualization at scale. Jesse Anderson is a phenomenal teacher who can explain - [Saying You Have Small Data Isn't Belittling Your Use Case](https://www.jesse-anderson.com/2018/07/saying-you-have-small-data-isnt-belittling-your-use-case/) - Blog Summary: (AI Summaries by Summarizes) Many engineers starting out with Big Data ask which technology to use for processing a dataset of 3 billion rows in 10,000 files that is 100 GB in size. The assumption is that small data technologies can’t handle this, but this is a misunderstanding of what Big Data is - [The Two Types of Data Engineering](https://www.jesse-anderson.com/2018/06/the-two-types-of-data-engineering/) - Blog Summary: (AI Summaries by Summarizes) There are two types of data engineering: SQL-focused and Big Data-focused. SQL-focused data engineering involves working with relational databases and processing data with SQL or a SQL-based language, sometimes using an ETL tool. Big Data-focused data engineering involves working with Big Data technologies like Hadoop, Cassandra, and HBase, and - [Why Real-time is the Future](https://www.jesse-anderson.com/2018/05/why-real-time-is-the-future/) - Blog Summary: (AI Summaries by Summarizes) Real-time Big Data is becoming increasingly important for organizations, teams, and individuals. However, in the past, we lacked the systems that could scale to the sizes and amounts of data needed for real-time processing. As a result, many organizations had to resort to batch processing over 24-hour windows, which - [The Four Types of Technologies You Need for Real-time Big Data Systems](https://www.jesse-anderson.com/2018/04/the-four-types-of-technologies-you-need-for-real-time-big-data-systems/) - Blog Summary: (AI Summaries by Summarizes) Real-time data pipelines bring new challenges and require new concepts and technologies to be learned and understood. Real-time data pipelines can be broken down into four general types: processors, analytics, ingestion and dissemination, and storage. Processors are responsible for processing incoming data and getting it ready for subsequent usage, - [What Are Batch and Real-time Big Data?](https://www.jesse-anderson.com/2018/04/what-are-batch-and-real-time-big-data/) - Blog Summary: (AI Summaries by Summarizes) Batch Big Data requires all data to be present before processing starts and can run over fixed periods of time, resulting in data being 1-2 times the time period old before it can be used. Common technologies used for batch processing in Big Data are Apache Hadoop and Apache - [Data engineers vs. data scientists](https://www.jesse-anderson.com/2018/04/data-engineers-vs-data-scientists/) - Blog Summary: (AI Summaries by Summarizes) The post discusses the differences between data engineers and data scientists. The author also talks about the role of machine learning engineers. The post can be found on the O’Reilly data blog. The author shares their latest thoughts and views on the topic. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”4.16″ global_colors_info=”{}”][et_pb_row admin_label=”row” - [Should You Even Do Big Data?](https://www.jesse-anderson.com/2018/03/should-you-even-do-big-data/) - Blog Summary: (AI Summaries by Summarizes) Big Data projects have a low success rate (usually 5-10%). Half-assed Big Data projects will fail without specific changes. Companies often reach out for help when their Big Data projects are failing. A cheap, quick, and easy fix is not possible for failing Big Data projects. The book “Data - [Unit Testing Kafka Streams](https://www.jesse-anderson.com/2018/03/unit-testing-kafka-streams/) - Blog Summary: (AI Summaries by Summarizes) Unit testing Kafka Streams code is important and can be done using the ProcessorTopologyTestDriver. To get started, you need to include the test libraries for Kafka Streams and Kafka in your Maven pom.xml file. Configuration properties are not needed for a unit test, but you can define topic names - [Are Your Programming Skills Ready for Big Data?](https://www.jesse-anderson.com/2018/02/are-your-programming-skills-ready-for-big-data/) - Blog Summary: (AI Summaries by Summarizes) Programming skills are crucial for working with Big Data. Programming skills can range from brand new to experienced in a language other than Java/Scala/Python. The level of programming skills needed depends on your role on the team. Big Data code is relatively small and self-contained, with the framework doing - [The Veteran Skill on a Data Engineering Team](https://www.jesse-anderson.com/2018/02/the-veteran-skill-on-a-data-engineering-team/) - Blog Summary: (AI Summaries by Summarizes) A veteran skill is often overlooked and unknown to data engineering teams. A project veteran is someone who has worked with Big Data and has had their solution in production for a while. A veteran brings a great deal of experience to the team and can prevent difficult code, - [How Much More Complicated Is Real-Time Big Data?](https://www.jesse-anderson.com/2018/01/how-much-more-complicated-is-real-time-big-data/) - Blog Summary: (AI Summaries by Summarizes) Big Data batch systems are 10x more complex than small data, while real-time systems are 15x more complex than small data. Programmers have the least increase in complexity when dealing with real-time systems, but they still need to learn the new system and any new APIs or concepts. Architects - [Getting Into Big Data as a Consultant](https://www.jesse-anderson.com/2018/01/getting-into-big-data-as-a-consultant/) - Blog Summary: (AI Summaries by Summarizes) Learning is the most important factor for success in Big Data consulting. Big Data is complex and requires advanced skills and knowledge. Certifications can be helpful, but proctored certifications are more valuable. Getting your first client requires a good reputation and trust. An awesome personal project can demonstrate your - [Q and A: How do I improve my skills to become a Data Engineer?](https://www.jesse-anderson.com/2017/12/q-and-a-how-do-i-improve-my-skills-to-become-a-data-engineer/) - Blog Summary: (AI Summaries by Summarizes) Programming is a crucial skill for data engineering, but it’s not the only one. System design and creation abilities are equally important. Understanding distributed systems and the technologies themselves is necessary to become a professional data engineer. Big infrastructure is required for data engineering, and it’s important to understand - [Why You Shouldn't Write Your Own Distributed System](https://www.jesse-anderson.com/2017/12/why-you-shouldnt-write-your-own-distributed-system/) - Blog Summary: (AI Summaries by Summarizes) Writing your own distributed system should not be taken lightly as it can have many ramifications. Poor reasons to write your own distributed system include not liking how a particular system works, thinking you can do it better, not having enough time to look for a system, or having - [What Is Big Data?](https://www.jesse-anderson.com/2017/11/what-is-big-data/) - Blog Summary: (AI Summaries by Summarizes) Big Data is defined as “can’t” – when a technical limitation prevents you from doing something. Examples of Big Data problems include being unable to add new features, run reports, or execute SQL statements due to technical limitations. This definition of Big Data is more helpful than the traditional - [You're Probably Not a Distributed Systems Engineer](https://www.jesse-anderson.com/2017/10/youre-probably-not-a-distributed-systems-engineer/) - Blog Summary: (AI Summaries by Summarizes) There are three main groups of teams that interact with distributed systems: users of end data products, users of existing distributed system frameworks, and creators of distributed systems frameworks. Users of end data products work with already created data pipelines and data products, while users of existing distributed system - [On Cheating with Big Data](https://www.jesse-anderson.com/2017/10/on-cheating-with-big-data/) - Blog Summary: (AI Summaries by Summarizes) To achieve the scales of Big Data, cheats or tradeoffs are necessary. HBase and Cassandra are both column-oriented NoSQL datastores, but their cheats are entirely different. HBase divides large tables into regions, while Cassandra divides them into partitions. HBase’s cheat of having one node serve all reads and writes - [When You Have the Wrong Team for Big Data](https://www.jesse-anderson.com/2017/09/when-you-have-the-wrong-team-for-big-data/) - Blog Summary: (AI Summaries by Summarizes) The right skills and people are crucial for the success of a Big Data project. Two teams were made up of the wrong skills and people, resulting in unsuccessful projects. The first team was a data warehousing team that was trained on Python programming and Big Data technologies. They - [Integration Testing for Kafka](https://www.jesse-anderson.com/2017/08/integration-testing-for-kafka/) - Blog Summary: (AI Summaries by Summarizes) Complex data pipelines and systems with Kafka are becoming more common, especially with microservices. Testing, debugging, and fixing these systems are crucial for ongoing success. Integration tests are longer running tests that test several different parts of a system, including microservices. Unit tests, on the other hand, test a - [How Training is Delivered From the Beginning to the End](https://www.jesse-anderson.com/2017/08/how-training-is-delivered-from-the-beginning-to-the-end/) - Blog Summary: (AI Summaries by Summarizes) Many training classes are useless because of bad decisions made during the budgeting phase. Budgeting for training is often arbitrary and not based on data, market price, or value. Going with a lower-priced training provider often results in poor curriculum and instructors. Curriculum ranges from free and terrible to - [This is Useless (Without Use Cases)](https://www.jesse-anderson.com/2017/07/this-is-useless-without-use-cases/) - Blog Summary: (AI Summaries by Summarizes) Understanding your use case is critical for the success of your project, especially in Big Data. Small data use cases can often use the same technology stack, while Big Data use cases require different technology stacks depending on the use case. Use case factors heavily into the design of - [Two Halves Don't Make a Whole](https://www.jesse-anderson.com/2017/07/two-halves-dont-make-a-whole/) - Blog Summary: (AI Summaries by Summarizes) Chapter 3 of the “Data Engineering Teams” book explains how to do a skill gap analysis. During the analysis, it’s a binary decision of whether a person has a skill or not. Some people want to use fractions to indicate partial skills, but this is a mistake. This thinking - [Apache Kafka and Amazon Kinesis](https://www.jesse-anderson.com/2017/07/apache-kafka-and-amazon-kinesis/) - Blog Summary: (AI Summaries by Summarizes) Apache Kafka and Amazon Kinesis are both contenders for Big Data messaging systems. Kinesis is a fully-managed streaming processing service available on AWS, while Kafka is an open source distributed publish subscribe system that can be installed on-premises or in the cloud. Kafka can store as much data as - [Moving from OmniGraffle to SVG](https://www.jesse-anderson.com/2017/07/moving-from-omnigraffle-to-svg/) - Blog Summary: (AI Summaries by Summarizes) OmniGraffle claims to support SVGs, but the import and export features do not work correctly. To move from OmniGraffle to SVG, you need to install Adobe Illustrator and copy the drawing from OmniGraffle to Illustrator, then save it as a compressed SVG file. To work with SVGZs, you need - [The Blame Game](https://www.jesse-anderson.com/2017/07/the-blame-game/) - Blog Summary: (AI Summaries by Summarizes) When a Big Data project fails, blame is often misplaced on the technology rather than the team itself. Teams need to look inwards at themselves to figure out why they’re failing, especially in Big Data due to its complexity. Management is responsible for the majority of failures in the - [What It Looks Like From the Outside](https://www.jesse-anderson.com/2017/06/what-it-looks-like-from-the-outside/) - Blog Summary: (AI Summaries by Summarizes) Big Data projects often fail due to incorrect assumptions made by management and engineering teams at the beginning of the project. Management may think that Hadoop/Spark/Big Data is a silver bullet or an easy rollout, which leads to problems later on. Teams often assume they will have time to - [Medium Data](https://www.jesse-anderson.com/2017/06/medium-data/) - Blog Summary: (AI Summaries by Summarizes) Many companies are experiencing a “witching hour” where their data is too big for small data and too small for Big Data, which is called “medium data.” Medium data is from companies that are looking to move from small data to Big Data, but don’t quite need the Big - [Beam 2.0 Q and A](https://www.jesse-anderson.com/2017/06/beam-2-0-q-and-a/) - Blog Summary: (AI Summaries by Summarizes) Apache Beam has just had its first API stable release, highlighting the growth of the project and increased usage in pre-production/development or production deployments. Beam is a cross-language and framework independent way of creating data pipelines, with a unified API for both batch and streaming computation. Beam is future-proof, - [The Difficulty of Transitioning to Data Pipelines](https://www.jesse-anderson.com/2017/06/the-difficulty-of-transitioning-to-data-pipelines/) - Blog Summary: (AI Summaries by Summarizes) Companies transitioning to Big Data, especially Kafka, face difficulty in moving from RPC-esque calls to a data pipeline where everything is exposed as raw data. Data pipelines are a new concept and have a loose coupling, unlike RPCs. Organizations need to answer questions related to socializing the data pipeline, - [Five Dysfunctions of a Data Engineering Team](https://www.jesse-anderson.com/2017/05/five-dysfunctions/) - Blog Summary: (AI Summaries by Summarizes) Companies are seeing efficiency gains and ROI from using Big Data technologies. However, the vast majority of teams fail and never get something into production. The top 5 reasons why data engineering teams fail at Big Data are discussed in a new talk based on the Data Engineering Teams - [How to Evaluate an Open Source Product](https://www.jesse-anderson.com/2017/05/how-to-evaluate-an-open-source-product/) - Blog Summary: (AI Summaries by Summarizes) Open source projects can be evaluated from a business point of view. When choosing between similar open source projects, consider the following: Companies that make money directly or indirectly off an open source project have a vested interest in maintaining and improving it. Open source projects that are running - [Kafka Topic Design Checklist](https://www.jesse-anderson.com/2017/05/kafka-topic-design-checklist/) - Blog Summary: (AI Summaries by Summarizes) Designing data for consumption in a Kafka topic requires more forethought than point-to-point consumption. When designing a Kafka topic, you need to decide on the name, schema, contents, key/ordering, number of partitions, and number of replicas. The topic name should be descriptive and not hardcoded in multiple places in - [The Many Meanings of Event-Driven Architecture: Kafka Edition](https://www.jesse-anderson.com/2017/05/the-many-meanings-of-event-driven-architecture-kafka-edition/) - Blog Summary: (AI Summaries by Summarizes) Kafka is often used to event changes and notify other systems of information changes or actions performed. The amount of information sent in an event should depend on how expensive a lookup is. If a lookup is expensive, event more information to save downstream lookups. An eventually consistent database - [Consumers and Creators of Technology](https://www.jesse-anderson.com/2017/05/consumers-and-creators-of-technology/) - Blog Summary: (AI Summaries by Summarizes) There are two types of companies: those that consume technology and those that create technology. Companies that consume technology use it to support or improve their non-technology product, while companies that create technology have their technology as their primary product. Companies that consume technology tend to place emphasis on - [There Are Several Hard Problems with Big Data](https://www.jesse-anderson.com/2017/04/there-are-several-hard-problems-with-big-data/) - Blog Summary: (AI Summaries by Summarizes) Big Data has several different hard problems that cannot be solved by changing just one thing. Big Data is 10-15x more complex than small data. The three main problems for Big Data are operations, development, and management. Management is crucial to the success of the project and problems tend - [Asking Better Questions](https://www.jesse-anderson.com/2017/04/asking-better-questions/) - Blog Summary: (AI Summaries by Summarizes) Asking good questions is an important life skill that applies to both personal and professional situations. Asking questions helps to clarify misunderstandings and gauge understanding of a topic. A good instructor can help draw out clarifying questions and recognize when a student should be asking a different question. When - [The Learning Disconnect](https://www.jesse-anderson.com/2017/04/the-learning-disconnect/) - Blog Summary: (AI Summaries by Summarizes) Educational material starts with a learning objective, which defines what the material is supposed to teach you. Before starting to learn anything, it’s important to be clear on your goal or desired outcome. If there is a mismatch between your goal and the learning objective, you won’t be successful. - [Doing Big Data ASAP](https://www.jesse-anderson.com/2017/04/doing-big-data-asap/) - Blog Summary: (AI Summaries by Summarizes) Apache Hive is a Big Data technology that uses a SQL-like language for its queries, reducing programming overhead to process data. To run Hive, you can turn to the cloud and use services like Amazon Web Services or Google Dataproc to spin up a cluster. Uploading data to the - [What Happens When You Hire a Data Scientist Without a Data Engineer](https://www.jesse-anderson.com/2017/03/what-happens-when-you-hire-a-data-scientist-without-a-data-engineer/) - Blog Summary: (AI Summaries by Summarizes) Data Scientists are often hired with the expectation that they will create models, but they may not have the necessary skills to create the data pipeline needed for those models. The definition of a Data Scientist is highly variable, and their programming and distributed system skill level can range - [Personal Project Data Sources](https://www.jesse-anderson.com/2017/03/personal-project-data-sources/) - Blog Summary: (AI Summaries by Summarizes) Having no previous Big Data experience is not a barrier to getting hired as a Data Engineer if you have a well-executed personal project that showcases your skills. Looking at available datasets can help you come up with an idea for your personal project and keep you focused. Some - [In-depth Interview](https://www.jesse-anderson.com/2017/03/in-depth-interview/) - Blog Summary: (AI Summaries by Summarizes) Justus Eapen interviewed Jesse Anderson for his podcast Hacker Practice. The interview covers a broad range of topics including creativity, Big Data, and Jesse’s book Data Engineering Teams. The interview delves deep into many data engineering concepts and theories. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”4.16″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”4.16″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” - [Kafka REST and JQuery Helper](https://www.jesse-anderson.com/2017/03/kafka-rest-and-jquery-helper/) - Blog Summary: (AI Summaries by Summarizes) The author is open sourcing a module they wrote for their Real-time Data Engineering class. The module uses Apache Spark and Apache Kafka to process data and display it in real-time on a webpage. The module is called KafkaRESTHelper and makes it easier to interface with the Kafka REST - [How Are Programming and Distributed Systems Different?](https://www.jesse-anderson.com/2017/02/how-are-programming-and-distributed-systems-different/) - Blog Summary: (AI Summaries by Summarizes) Programming and distributed systems are two different skills needed in a data engineering team. Programming can be divided into three types of programmers: coders, simple programmers, and advanced programmers. Distributed systems are not easy to work with, and Big Data frameworks only make it easier to concentrate on the - [Announcement: Data Engineering Teams Book](https://www.jesse-anderson.com/2017/02/announcement-data-engineering-teams-book/) - Blog Summary: (AI Summaries by Summarizes) 85% of Big Data projects fail to make it into production, according to Gartner’s research. The author has written a book on Big Data, sharing the secrets to make Big Data projects succeed. The book is primarily written for managers, VPs, and CxOs who are managing teams creating a - [Is Kafka Only a Big Data Tool?](https://www.jesse-anderson.com/2017/02/is-kafka-only-a-big-data-tool/) - Blog Summary: (AI Summaries by Summarizes) Kafka is not just a Big Data tool and can be used for small data as well. For most Big Data technologies, not having or having a Big Data problem in the future is the reason not to use them. Kafka is a distributed publish subscribe system that can - [What Do I Look for in Data Engineers?](https://www.jesse-anderson.com/2017/02/what-do-i-look-for-in-data-engineers/) - Blog Summary: (AI Summaries by Summarizes) A strong programming background is crucial for Data Engineers, with many having a Master’s degree or above in Computer Science with a focus on distributed systems or data. The best Data Engineers are not content with just programming and have started to cross-train into other fields, such as data - [Q and A: Ingesting into Hadoop](https://www.jesse-anderson.com/2017/01/q-and-a-ingesting-into-hadoop/) - Blog Summary: (AI Summaries by Summarizes) Apache Sqoop is a tool that can move data from a RDBMS and put it into HDFS or HBase, and vice versa. There are several ways to do simple file transfers into HDFS, including using Apache Oozie, Hue’s REST interface, Hadoop’s WebHDFS REST or FUSE interfaces, or writing a - [Hadoop MapReduce Dedupe Algorithm](https://www.jesse-anderson.com/2017/01/hadoop-mapreduce-dedupe-algorithm/) - Blog Summary: (AI Summaries by Summarizes) The video demonstrates live coding of a dedupe algorithm. The algorithm is used to remove duplicates from several data files. The video shows a simple version of the algorithm and a more complicated version with custom logic. The video is a helpful resource for those interested in learning how - [How much do companies lose before training?](https://www.jesse-anderson.com/2017/01/how-much-do-companies-lose-before-training/) - Blog Summary: (AI Summaries by Summarizes) Starting to write code or design a solution before receiving proper training is a bad idea, especially in Big Data. Making a mistake with small data isn’t costly and can be fixed quickly, but making a mistake with Big Data is very costly and can take a while to - [Maven Tips](https://www.jesse-anderson.com/2017/01/maven-tips/) - Blog Summary: (AI Summaries by Summarizes) Working with complex and multi-module Maven projects can be challenging. To list all modules in a project, run the command `mvn help:all-profiles`. The output will show the project’s reactor build order and list all profiles for each module. The command can be helpful for navigating and understanding the structure - [Apache Beam Regex](https://www.jesse-anderson.com/2016/12/apache-beam-regex/) - Blog Summary: (AI Summaries by Summarizes) The Regex class in Beam allows for distributed string processing. The interface of the Regex class is designed to be familiar to Java developers. The Regex.find() method can be used to filter a file based on a regular expression. The Regex.find() method can also be used to extract specific - [Beam's Pico WordCount](https://www.jesse-anderson.com/2016/12/beams-pico-wordcount/) - Blog Summary: (AI Summaries by Summarizes) The goal of the game in Big Data frameworks is to create the fewest lines of code for WordCount. The author is a committer on Apache Beam and has recently improved the regular expression handling in Beam. The smallest WordCount using Beam involves reading in a file, using the - [Are you attaining your goals?](https://www.jesse-anderson.com/2016/11/are-you-attaining-your-goals/) - Blog Summary: (AI Summaries by Summarizes) Reflect on how you did this year before making goals for the next year. Becoming a Data Engineer requires a significant amount of time, effort, and skill. Check if you have achieved your goal of becoming a Data Engineer. If you’re not making progress towards your goal, you need - [Unit Testing Kafka Consumers](https://www.jesse-anderson.com/2016/11/unit-testing-kafka-consumers/) - Blog Summary: (AI Summaries by Summarizes) Unit testing your Kafka code is crucial, especially for your Consumers. Refactor your Consumer code to be able to change it at runtime and create a separate method for creating the KafkaConsumer. Refactor the code that consumes data from the Consumer object to be callable from the unit test - [Unit Testing Kafka](https://www.jesse-anderson.com/2016/11/unit-testing-kafka/) - Blog Summary: (AI Summaries by Summarizes) Unit testing Kafka code is crucial as it transports important data. As of version 0.9.0, there is a new way to unit test with mock objects. To refactor the producer, change the `Producer` to an interface and create a separate method for creating the `KafkaProducer`. Refactor the code that - [What will become of Big Data?](https://www.jesse-anderson.com/2016/11/what-will-become-of-big-data/) - Blog Summary: (AI Summaries by Summarizes) Big Data technologies will continue to mature over the next 5-10 years. Better stories on the things enterprises need will emerge. Technologies for metadata management and granular authorization will improve. Hadoop MapReduce will gradually phase out, while Apache Spark will mature and stabilize. Data Engineers will face difficulty in - [Why should or shouldn't you become a Data Engineer?](https://www.jesse-anderson.com/2016/10/why-should-or-shouldnt-you-become-a-data-engineer/) - Blog Summary: (AI Summaries by Summarizes) Becoming a Data Engineer can have significant financial benefits, with potential increases of $20,000 to $60,000 per year compared to other data-related positions. There is currently a high demand and low supply of qualified Data Engineers, making it a lucrative field to pursue. Data is changing the way businesses - [Q and A: How can I tear out Informatica and MySQL and put in Big Data?](https://www.jesse-anderson.com/2016/10/q-and-a-how-can-i-tear-informatica-and-mysql-and-put-in-big-data/) - Blog Summary: (AI Summaries by Summarizes) The post addresses three questions from a subscriber on gaining a hands-on understanding of Big Data technologies, convincing the engineering team to change directions, and tearing Informatica and MySQL away from the engineering team. The author created a course specifically for management called the Business of Big Data, which - [Getting Stuck Crawling with Big Data](https://www.jesse-anderson.com/2016/10/getting-stuck-crawling-with-big-data/) - Blog Summary: (AI Summaries by Summarizes) Breaking down Big Data projects into smaller pieces is recommended, using the crawl, walk, run process. Some companies get stuck at the crawl phase and don’t progress to the walk and run phases. Stopping at crawl looks like cloning the data warehouse in Hadoop without improving or using new - [Strata+Hadoop World and Trends](https://www.jesse-anderson.com/2016/10/stratahadoop-world-and-trends/) - Blog Summary: (AI Summaries by Summarizes) Strata+Hadoop World is the Super Bowl of Big Data conferences where the best minds talk about the present and future conditions of Big Data. The first session covered Apache Beam and some of the interesting features for Big Data. The second session covered how Apache Spark and Java can - [Big Data Minority Scholarships](https://www.jesse-anderson.com/2016/09/big-data-minority-scholarships/) - Blog Summary: (AI Summaries by Summarizes) Lack of diversity in technology, especially in specialties like Big Data, is a problem of supply and not demand. Teams benefit from diversity, which comes in various forms, from ethnicity to socioeconomic backgrounds. Diverse teams don’t suffer from groupthink and single point of view. A scholarship for online data - [Solving the First and Last Mile Problem With Kafka Part 2](https://www.jesse-anderson.com/2016/09/solving-the-first-and-last-mile-problem-with-kafka-part-2/) - Blog Summary: (AI Summaries by Summarizes) Big Data faces challenges in finding value in large amounts of data, especially in real-time systems. Kafka provides several windows into real-time streams, including a REST proxy for creating real-time dashboards. Before creating a dashboard, data must be processed and analyzed upstream, and the dashboard should consume processed data - [Solving the First and Last Mile Problem With Kafka Part 1](https://www.jesse-anderson.com/2016/09/solving-the-first-and-last-mile-problem-with-kafka-part-1/) - Blog Summary: (AI Summaries by Summarizes) Apache Kafka is a great technology for moving, storing, and processing data in real-time. Before Kafka, Apache Flume was one of the most commonly used technologies for moving data in real-time, but it lacked a great place to place the data. Kafka eliminates the problems of the first and - [How Programmers Should Start Viewing Training](https://www.jesse-anderson.com/2016/09/how-programmers-should-start-viewing-training/) - Blog Summary: (AI Summaries by Summarizes) Programmers don’t invest in themselves like other professions do. This limits their growth and career advancement compared to business and marketing professionals. Developers often expect their companies to train them, but this mindset can lead to falling behind on the latest technology trends. Investing in oneself through high-quality training - [On complexity in big data](https://www.jesse-anderson.com/2016/09/on-complexity-in-big-data/) - Blog Summary: (AI Summaries by Summarizes) The author has written an in-depth piece on why Big Data isn’t easy, cheap, or quick. The piece was published on O’Reilly. The author has years of experience teaching Big Data. [et_pb_section admin_label=”section”] [et_pb_row admin_label=”row”] [et_pb_column type=”4_4″][et_pb_text admin_label=”Text”] After years of teaching Big Data, I’ve come up with the - [Q and A: Is a Data Engineer the same thing as a BI or DBA?](https://www.jesse-anderson.com/2016/08/q-and-a-is-a-data-engineer-the-same-thing-as-a-bi-or-dba/) - Blog Summary: (AI Summaries by Summarizes) A Data Engineer is someone who specializes in creating software solutions around data, predominantly based around Hadoop, Spark, and the open source Big Data ecosystem. Data Engineers are not the same as DBAs, Business Intelligence, Data Analysts, or ETL Developers, but people with these titles can become Data Engineers - [Crawl, Walk, Run with Big Data](https://www.jesse-anderson.com/2016/08/crawl-walk-run-with-big-data/) - Blog Summary: (AI Summaries by Summarizes) Attacking a Big Data project with an all-or-nothing mindset leads to failure Breaking the project into manageable phases called crawl, walk, and run is recommended Crawling phase involves doing the minimum to start using Big Data, such as getting current data into Hadoop and setting up systems to bring - [Q and A: Big Data strategy](https://www.jesse-anderson.com/2016/08/q-and-a-big-data-strategy/) - Blog Summary: (AI Summaries by Summarizes) Big Data strategy requires a different mindset than small data strategy. Lack of understanding of Big Data technologies can lead to failure in a Big Data project. A Big Data project requires qualified Data Engineers with technical understanding. Both the manager and engineers need to attend relevant classes to - [Is Big Data Cheap?](https://www.jesse-anderson.com/2016/08/is-big-data-cheap/) - Blog Summary: (AI Summaries by Summarizes) Big Data is not always cheap, some things are cheap and some things are more expensive. Hadoop is the gold standard for both startups and enterprises, and it is not an open source knock off of a better closed source framework. Small data solutions often need a single computer, - [Apache Kafka and Google Cloud Pub/Sub](https://www.jesse-anderson.com/2016/07/apache-kafka-and-google-cloud-pubsub/) - Blog Summary: (AI Summaries by Summarizes) Apache Kafka, Google Cloud Pub/Sub, and Amazon Kinesis are contenders for Big Data messaging systems. Pub/Sub is a cloud service provided by Google Cloud, while Kafka can be installed on-premises or in the cloud. Kafka can store as much data as needed and supports log compaction, while Pub/Sub stores - [Kafka 0.10 Changes for Developers](https://www.jesse-anderson.com/2016/07/kafka-0-10-changes-for-developers/) - Blog Summary: (AI Summaries by Summarizes) Kafka 0.10 is out with some changes that developers need to know about. The KafkaConsumer now allows you to specify a maximum number of messages to return by using the max.poll.records property to a number in the KafkaConsumer. The major change for this version is the addition of Kafka - [Question and Answers with the Apache Beam Team](https://www.jesse-anderson.com/2016/07/question-and-answers-with-the-apache-beam-team/) - Blog Summary: (AI Summaries by Summarizes) Apache Beam is a unified programming model for creating batch and stream data processing pipelines. The project just had its first release and is now working towards the second release, 0.2.0-incubating. Interviewees are a mix of committers and users of Beam. Interviewees describe Beam as an abstraction for stream - [Ability Gap - Why We Need Data Engineers](https://www.jesse-anderson.com/2016/07/ability-gap-why-we-need-data-engineers/) - Blog Summary: (AI Summaries by Summarizes) Big Data is a complicated field with many new technologies and changes within technologies that make it time prohibitive to keep up with. There is an ability gap when it comes to Big Data concepts, where some people simply won’t understand them on their best day. The level of - [Big Data's Required and Recommended Technical Skills](https://www.jesse-anderson.com/2016/06/big-datas-required-and-recommended-technical-skills/) - Blog Summary: (AI Summaries by Summarizes) Hadoop requires technical skills to get started with big data. For developers, required skills include intermediate to advanced Java knowledge and a general understanding of Linux. Recommended skills for developers include knowledge of Scala and SQL, as well as a background in distributed systems. For administrators, required skills include - [My Big Data Journey](https://www.jesse-anderson.com/2016/06/my-big-data-journey/) - Blog Summary: (AI Summaries by Summarizes) The author’s background in creating distributed systems from scratch in Java gave them a leg up in learning Big Data technologies. Multi-threading experience was also helpful in understanding how to move data around between threads and processes in Big Data frameworks. The author worked at a mobile startup that - [The Case for Heron](https://www.jesse-anderson.com/2016/06/the-case-for-heron/) - Blog Summary: (AI Summaries by Summarizes) Twitter has open sourced Heron, which replaced Storm at Twitter and has been running in production for over 2 years. Heron continues to use Storm’s API, allowing companies with large Storm codebases to get performance improvements without having to rewrite everything. Heron provides a 2-5x performance boost without changing - [We Live, Eat, and Breathe This Stuff](https://www.jesse-anderson.com/2016/05/we-live-eat-and-breathe-this-stuff/) - Blog Summary: (AI Summaries by Summarizes) Learning big data and its expanding ecosystem takes time and effort, even for those with a background in distributed systems. There are many projects and nuances to learn, making it difficult for newcomers to Big Data to understand the full scope and potential of the technology. When integrating with - [Spark and Java - Yes, They Work Together](https://www.jesse-anderson.com/2016/05/spark-and-java-yes-they-work-together/) - Blog Summary: (AI Summaries by Summarizes) Learning two new and different technologies at the same time makes you catch neither. Java programmers come, see the Java API with Spark, and decide to learn Scala. The Scala code is just more concise and readable. Spark’s Java API with Java 7 eyes and saw Java 7 code - [SSH With Google Cloud](https://www.jesse-anderson.com/2016/04/ssh-with-google-cloud/) - Blog Summary: (AI Summaries by Summarizes) To SSH into your Google Cloud instance, you need to create a new SSH key using the `ssh-keygen` program. The program will create a public and private key pair for you to use with Google Cloud. To get the public key for your Google Cloud Console, run `cat ~/.ssh/google_compute_engine.pub` - [Announcement: Creating Big Data Solutions with Impala](https://www.jesse-anderson.com/2016/04/creating-big-data-solutions-with-impala/) - Blog Summary: (AI Summaries by Summarizes) The author’s latest screencast on Apache Impala called “Creating Big Data Solutions with Impala” was released on O’Reilly. The author’s relationship with Impala started when he joined Cloudera and offered to help with various projects, including Impala. The author created Cloudera’s first public VM image to help people use - [Unit Testing Spark with Java](https://www.jesse-anderson.com/2016/04/unit-testing-spark-with-java/) - Blog Summary: (AI Summaries by Summarizes) Unit testing, Apache Spark, and Java can work well together. Unit testing is important for Big Data code to ensure faster turnaround time on fixes and working on code. Holden Karau released Spark Testing Base, a Spark unit testing framework, at Strata NYC 2015. Refactoring code is necessary to - [Three Top Themes From Strata+Hadoop World](https://www.jesse-anderson.com/2016/04/three-top-themes-from-stratahadoop-world/) - Blog Summary: (AI Summaries by Summarizes) Real-time Big Data is becoming increasingly popular and companies are finding that it gives them an advantage and agility they didn’t have before. Real-time systems like Kafka allow Data Scientists to get from hypothesis to production quicker and run and score using several models at the same time. Second, - [Kafka 0.9.0 Changes For Developers](https://www.jesse-anderson.com/2016/04/kafka-0-9-0-changes-for-developers/) - Blog Summary: (AI Summaries by Summarizes) Kafka 0.9.0 has new features for developers The most notable change is a brand-new consumer API The new consumer is a great replacement for the SimpleConsumer and the old consumer The JavaDocs for the new consumer have all of the sample code The new consumer can easily consume a - [The ROI of the Right Training](https://www.jesse-anderson.com/2016/03/the-roi-of-the-right-training/) - Blog Summary: (AI Summaries by Summarizes) Investing in knowledge provides the highest ROI. Training can save individuals and teams significant amounts of time. The break-even point for training costs is typically around 2-4 weeks of time saved. Great instructors can provide valuable insights and help avoid mistakes. The intangible benefits of training, such as avoiding - [Hadoop Cheat Sheet](https://www.jesse-anderson.com/2016/03/hadoop-cheat-sheet/) - Blog Summary: (AI Summaries by Summarizes) Hadoop has a large developer community with projects that have names that don’t correlate to their function. This cheat sheet helps keep track of the different projects in the Hadoop ecosystem and their respective functions. The projects are broken up into three categories: Distributed Systems, Processing Data, and Getting - [Identifying Great Training Even If You Know Nothing About the Subject](https://www.jesse-anderson.com/2016/03/identifying-great-training-even-if-you-know-nothing-about-the-subject/) - Blog Summary: (AI Summaries by Summarizes) Identifying the end goal is the first step in identifying great training. Look at job listings for your desired end goal to identify the common technologies and qualifications needed. Evaluating classes involves looking at the class itself and the reviews of the class. Reviews should be taken with a - [The Hidden Costs of the Wrong Training](https://www.jesse-anderson.com/2016/03/the-hidden-costs-of-training/) - Blog Summary: (AI Summaries by Summarizes) The biggest cost for training is your time, which is also known as opportunity cost. Good training can provide the best ROI, but bad training can cost more than the amount spent. For individuals, the cost breakdown for a course includes the course cost, lecture cost, and practice cost, - [Is My Developer Team Ready for Big Data?](https://www.jesse-anderson.com/2016/03/is-my-developer-team-ready-for-big-data/) - Blog Summary: (AI Summaries by Summarizes) The most common question asked by business leaders is whether their developer team is ready for big data projects. Executives understand the potential benefits of big data projects but are unsure if their current development teams have the necessary skills to create solutions. The full guest post can be - [Why You're Not Getting a Data Engineering Job](https://www.jesse-anderson.com/2016/03/why-youre-not-getting-a-data-engineering-job/) - Blog Summary: (AI Summaries by Summarizes) A hiring manager for a large company looking for Data Engineers passed on job candidates because they weren’t qualified enough. The interviewees may have bought a terrible online or in-person course that made unjustified promises. The problem is not that the hiring manager wouldn’t accept people who went through - [The Cloudera Experience](https://www.jesse-anderson.com/2014/12/the-cloudera-experience/) - Blog Summary: (AI Summaries by Summarizes) Big Data and Hadoop companies are expected to see a sharp uptick, starting with Hortonworks’ recent IPO. Cloudera, a notable absence in the IPO chase, is characterized by its incredibly smart, humble, and dedicated employees. Cloudera employees are assumed to have a high level of technical competence, and former - [HBase With Playing Cards](https://www.jesse-anderson.com/2014/08/hbase-with-playing-cards/) - Blog Summary: (AI Summaries by Summarizes) A 7-minute video demonstrating how HBase works with playing cards has been created. The video can be accessed through the embedded YouTube link. [et_pb_section admin_label=”section”] [et_pb_row admin_label=”row”] [et_pb_column type=”4_4″][et_pb_text admin_label=”Text”]I created a 7 minute video showing how HBase works with playing cards. [/et_pb_text][/et_pb_column] [/et_pb_row] [/et_pb_section] - [More NFL Posts](https://www.jesse-anderson.com/2014/01/more-nfl-posts/) - Blog Summary: (AI Summaries by Summarizes) Two guest posts have been published about the NFL Play-by-Play project. The first post is on GigaOM and provides insights for business leaders to learn from football. The second post is on Cloudera’s VISION blog and offers a more technical explanation of the project. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row - [Hiring Your First Software Engineer](https://www.jesse-anderson.com/2013/12/hiring-your-first-software-engineer/) - Blog Summary: (AI Summaries by Summarizes) An article on CEO.com was posted discussing ways to hire and interview your first software engineer. The article emphasizes the importance of giving back to software groups as you use their help. The author also wrote a guest post for startup communities discussing how to give back. [et_pb_section fb_built=”1″ - [Processing Big Data with MapReduce](https://www.jesse-anderson.com/2013/07/processing-big-data-with-mapreduce/) - Blog Summary: (AI Summaries by Summarizes) Jesse Anderson has released a new series of screencasts on Hadoop MapReduce, published by Pragmatic Programmers. The screencasts are a great way for beginners to learn about Hadoop. A virtual machine is available with everything needed to run Hadoop, MapReduce, and Eclipse. Source code for the screencasts is available - [Cloudera QuickStart VM and Eclipse](https://www.jesse-anderson.com/2013/07/cloudera-quickstart-vm-and-eclipse/) - Blog Summary: (AI Summaries by Summarizes) You can create a MapReduce project in Eclipse and debug it. The Cloudera QuickStart VM lets developers get started with writing MapReduce code without having to worry about software installs and configuration. Eclipse is installed on the VM and there is a link on the desktop to start it. - [Guest Post on O'Reilly Programming Blog](https://www.jesse-anderson.com/2013/07/guest-post-on-oreilly-programming-blog/) - Blog Summary: (AI Summaries by Summarizes) The author has written a guest post for the O’Reilly Programming Blog. The post discusses augmenting datasets and dealing with unstructured data. The post is related to the author’s upcoming talk at OSCON. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text - [HBase REST Interface Part 3](https://www.jesse-anderson.com/2013/07/hbase-rest-interface-part-3/) - Blog Summary: (AI Summaries by Summarizes) The blog series on REST has ended with the last post covering getting rows using JSON and XML. The post includes full code samples with comments. The final post in the series can be found on the Cloudera blog. The code samples can be accessed on GitHub. [et_pb_section fb_built=”1″ - [Propose, Prepare, Present Review](https://www.jesse-anderson.com/2013/06/propose-prepare-present-review/) - Blog Summary: (AI Summaries by Summarizes) The book “Propose, Prepare, Present” by Alistair Croll is a great resource for improving conference submissions. The book covers not only the tactics of what makes a submission good or bad, but also the behind-the-scenes of the conference industry. The book shows some of the cardinal sins one can - [Million Monkeys Interview on The Creators Project](https://www.jesse-anderson.com/2013/05/million-monkeys-interview-on-the-creators-project/) - Blog Summary: (AI Summaries by Summarizes) The post is an interview about The Million Monkeys Project. The interview was conducted for The Creators Project. The interview discusses other creative uses of Shakespeare’s works. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text admin_label=”Text” _builder_version=”4.11.4″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” hover_enabled=”0″ - [Interview in the RGJ](https://www.jesse-anderson.com/2013/05/interview-in-the-rgj/) - Blog Summary: (AI Summaries by Summarizes) The post is an interview with the author for the Reno Gazette Journal. The author is seen coding barefoot in the accompanying video. The link to the interview is provided in the post. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text - [HBase REST Interface Part 2](https://www.jesse-anderson.com/2013/04/hbase-rest-interface-part-2/) - Blog Summary: (AI Summaries by Summarizes) The blog post covers adding rows using JSON and XML in HBase REST interface. The author has written a series of blogs on HBase REST interface. The post provides full code samples with comments. The second post in the series is available on the Cloudera blog. The author has - [HBase REST Interface](https://www.jesse-anderson.com/2013/03/hbase-rest-interface/) - Blog Summary: (AI Summaries by Summarizes) The HBase REST interface allows non-Java programmers to access HBase over a REST interface. The author wrote a series of blogs on the HBase REST interface. The first post in the series is up on the Cloudera blog. Full code samples with comments are available on GitHub. [et_pb_section fb_built=”1″ - [Startup Communities Guest Post](https://www.jesse-anderson.com/2013/03/startup-communities-guest-post/) - Blog Summary: (AI Summaries by Summarizes) The guest post written by the assistant for Brad Feld’s Startup Communities was published. The post discusses the role of feeder groups like NNSDG in helping entrepreneurs. The post is titled “I’ve Been a Bad Feeder”. The post can be found at http://www.startuprev.com/how-do-tech-groups-act-as-community-feeders/. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” - [Optimizing MapReduce Jobs](https://www.jesse-anderson.com/2013/02/optimizing-mapreduce-jobs/) - Blog Summary: (AI Summaries by Summarizes) The post is about improving MapReduce performance. The post can be found on the Cloudera blog. The post includes graphs and code to help readers try out the suggestions. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text admin_label=”Text” _builder_version=”4.11.4″ background_size=”initial” background_position=”top_left” - [Social Recruiting The Best](https://www.jesse-anderson.com/2013/01/social-recruiting-the-best/) - Blog Summary: (AI Summaries by Summarizes) Cloudera is hiring and using social recruiting to find the best candidates. The job post is using OnGig, which provides all the information about the position and answers to common questions on the page. OnGig also includes a video about the company and why someone should work there. This - [MapReduce and Graph Theory](https://www.jesse-anderson.com/2013/01/mapreduce-and-graph-theory/) - Blog Summary: (AI Summaries by Summarizes) The post discusses using MapReduce and Graph Theory to solve a Boggle roll. The post includes an explanation and code for readers to try out. The post can be found on the Cloudera blog. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” - [Mary Alice Berg (1936-2013)](https://www.jesse-anderson.com/2013/01/mary-alice-berg-1936-2013/) - Blog Summary: (AI Summaries by Summarizes) Mary Alice Berg was the author’s grandmother who passed away at the age of 76 due to renal cancer. She was a Jehovah’s Witness and they don’t believe in eulogizing a person after they die. The author remembers her for the good times spent with her, not the things - [Big Data Congress](https://www.jesse-anderson.com/2013/01/big-data-congress/) - Blog Summary: (AI Summaries by Summarizes) The author will be speaking at the Big Data Congress in St. John, New Brunswick, Canada. The session’s theme is the software that powers Big Data. The author will be sharing Million Monkeys stories and answering questions. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ - [NFL Play By Play Analysis](https://www.jesse-anderson.com/2013/01/nfl-play-by-play-analysis/) - Blog Summary: (AI Summaries by Summarizes) Advanced NFL Stats has released the play by play data of the 2002 season. The author performed a quick analysis of the data using Hive and MapReduce to look at incomplete passes. The code used for the analysis is available on the author’s GitHub account. The author created two - [Articles In Pragmatic Magazine](https://www.jesse-anderson.com/2012/12/articles-in-pragmatic-magazine/) - Blog Summary: (AI Summaries by Summarizes) Two articles have been published in Pragmatic Magazine to help with moving to The Cloud. The first article discusses the politics of moving to The Cloud in a company. The second article provides insights into how The Cloud can save money. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” - [The Screencast Experience](https://www.jesse-anderson.com/2012/11/the-screencast-experience/) - Blog Summary: (AI Summaries by Summarizes) The author created a screencast that was published by Pragmatic Programmers. Before starting the project, the author couldn’t find any blog posts about the experience of creating a screencast. Creating a screencast is like writing a novella, creating custom graphics per slide, and doing the audiobook for the novella - [The Cloud and Amazon Web Services](https://www.jesse-anderson.com/2012/10/the-cloud-and-amazon-web-services/) - Blog Summary: (AI Summaries by Summarizes) Jesse Liberty has created a series of screencasts on The Cloud and Amazon Web Services, published by Pragmatic Programmers. The series starts with basic concepts behind cloud computing and guides users through practical, hands-on examples using Amazon’s cloud offerings. The series targets users who are new to the cloud, - [Kinder Gentler Apple](https://www.jesse-anderson.com/2012/09/kinder-gentler-apple/) - Blog Summary: (AI Summaries by Summarizes) Tim Cook took over as CEO of Apple after Steve Jobs. Steve Jobs had a vendetta against Flash, but it was not necessarily based on technical reasons. Other products, like Unity, were doing similar things to Flash but were exempted from Apple’s App Store process. There is no indication - [Strata + Hadoop World 2012](https://www.jesse-anderson.com/2012/09/strata-hadoop-world-2012/) - Blog Summary: (AI Summaries by Summarizes) The author will be speaking at the Strata and Hadoop World conference. The topic of the author’s talk will be the Million Monkeys project, randomness, and using Hadoop in business. Attending the conference is a great way to learn more about Hadoop. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ - [Hadoop The Definitive Guide 3rd Edition Review](https://www.jesse-anderson.com/2012/08/hadoop-the-definitive-guide-3rd-edition-review/) - Blog Summary: (AI Summaries by Summarizes) The 3rd edition of “Hadoop The Definitive Guide” covers the latest changes to the 1.x and 2.x APIs. The book extensively discusses the new distributed resource management system named YARN. The new edition also covers the new features of HDFS, including high availability and federation. “Hadoop The Definitive Guide” - [A Little Too Quiet](https://www.jesse-anderson.com/2012/07/a-little-too-quiet/) - Blog Summary: (AI Summaries by Summarizes) The author has been busy and things have been quiet on their blog. The author has changed jobs and is now a Curriculum Developer and Instructor at Cloudera. The author will be teaching more about Hadoop and its ecosystem. The author has been creating new content, some of which - [EC2 Performance, Spot Instance ROI and EMR Scalability](https://www.jesse-anderson.com/2012/02/ec2-performance-spot-instance-roi-and-emr-scalability/) - Blog Summary: (AI Summaries by Summarizes) Amazon introduced Elastic Compute Cloud (EC2) to Amazon Web Services (AWS) in 2006 and Elastic MapReduce (EMR) in 2009. EMR uses Hadoop to create MapReduce jobs using EC2 instances with Simple Storage Service (S3) as the permanent storage mechanism. Spot Instances allow you to bid on EMR or EC2 - [Million Monkeys Visualization](https://www.jesse-anderson.com/2011/10/million-monkeys-visualization/) - Blog Summary: (AI Summaries by Summarizes) A new visualization of the Million Monkeys’ data was created at the recent Hack4Reno event. The visualization allows users to select their favorite work of Shakespeare and find out how many times a particular character was found. The data was generated from ~3GB of raw monkey data and converted - [A Few Million Monkeys Randomly Recreate Every Work Of Shakespeare](https://www.jesse-anderson.com/2011/10/a-few-million-monkeys-randomly-recreate-every-work-of-shakespeare/) - Blog Summary: (AI Summaries by Summarizes) The Million Monkeys project successfully recreated all 38 works of Shakespeare through random generation of character groups. The project went viral on September 26, 2011, with over 25,000 unique visitors and 300 sites referring traffic from 119 countries. The project was open-sourced and over 7.5 trillion character groups were - [A Few Million Monkeys Randomly Recreate Shakespeare](https://www.jesse-anderson.com/2011/09/a-few-million-monkeys-randomly-recreate-shakespeare/) - Blog Summary: (AI Summaries by Summarizes) A group of virtual monkeys has successfully recreated every work of Shakespeare through random gibberish. The project started on August 21, 2011, and over 6.5 trillion character groups have been randomly generated and checked out of the 5.5 trillion possible combinations. The monkeys will continue typing until every work - [A Few More Million Amazonian Monkeys](https://www.jesse-anderson.com/2011/08/a-few-more-million-amazonian-monkeys/) - Blog Summary: (AI Summaries by Summarizes) The project aims to recreate every work of Shakespeare randomly using virtual, computerized monkeys that output random gibberish. The computer program compares the monkey’s gibberish to every work of Shakespeare to see if it matches a small portion of what Shakespeare wrote. The monkeys’ data from Amazon’s cloud is - [Pitching Agile](https://www.jesse-anderson.com/2011/07/pitching-agile/) - Blog Summary: (AI Summaries by Summarizes) Before implementing Agile Software Development methodology, it is important to get buy-in and support from as many people as possible. This presentation provides guidance on how to formulate an argument for Agile based on the person’s department or position. The presentation is divided into three parts, each with a - [Post Agile Checklist](https://www.jesse-anderson.com/2011/07/post-agile-checklist/) - Blog Summary: (AI Summaries by Summarizes) The presentation discusses ways to fully utilize Agile Software Development methodology in teams or companies. The presentation is divided into two parts, each with an embedded YouTube video. The presentation also includes a post-agile checklist, which is embedded as a slideshow from Slideshare. The checklist provides a list of - [Oracle Database 9i, 10g, and 11g Programming Techniques and Solutions Review](https://www.jesse-anderson.com/2011/07/oracle-database-9i-10g-and-11g-programming-techniques-and-solutions-review/) - Blog Summary: (AI Summaries by Summarizes) The book “Expert Oracle Database Architecture: Oracle Database 9i, 10g, and 11g Programming Techniques and Solutions, Second Edition” by Tom Kyte is highly recommended for those who want to learn the deep inner workings of Oracle and get copious information on the topics. The target audience is not a - [RIM Employee's E-Mail To CEO](https://www.jesse-anderson.com/2011/06/rim-employees-e-mail-to-ceo/) - Blog Summary: (AI Summaries by Summarizes) A Research in Motion (RIM) employee wrote an email to the CEO addressing the problems at RIM and why the company is on a downward spiral. The email could apply to a lot of technology companies. RIM responded to the email with a swing and miss. More employees chimed - [Microsoft Kinect SDK](https://www.jesse-anderson.com/2011/06/microsoft-kinect-sdk/) - Blog Summary: (AI Summaries by Summarizes) Microsoft has released an SDK for the Kinect. The SDK provides good documentation and a video quickstart. Programming with the Kinect is now easier with the SDK. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text admin_label=”Text” _builder_version=”3.27.4″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”] - [A Million Amazonian Monkeys](https://www.jesse-anderson.com/2011/06/a-million-amazonian-monkeys/) - Blog Summary: (AI Summaries by Summarizes) The author created a project using Hadoop’s MapReduce with Amazon’s Elastic MapReduce to simulate the Infinite Monkey Theorem. The project uses fake Amazonian Map Monkeys to create random data in ASCII between a and z, which is then passed through a Bloom Field membership test and a string comparison - [Googler Waves Goodbye](https://www.jesse-anderson.com/2011/06/googler-waves-goodbye/) - Blog Summary: (AI Summaries by Summarizes) A member of Google’s Wave team left Google and explained why in a blog post. The member stated that “Google’s vaunted scalable software infrastructure is obsolete.” No explanation or proof was given for this statement, but it was suggested that the software stack on top of the infrastructure is - [Opinions on Interviews](https://www.jesse-anderson.com/2011/05/opinions-on-interviews/) - Blog Summary: (AI Summaries by Summarizes) Software interviews are broken. A good write-up on other suggestions for interviews can be found at http://techcrunch.com/2011/05/07/why-the-new-guy-cant-code/. The best suggestion is to not interview anyone who hasn’t accomplished anything. Certificates and degrees are not accomplishments. Real-world projects with real-world users are considered as accomplishments. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row - [Apple's Hardware and Software](https://www.jesse-anderson.com/2011/05/apples-hardware-and-software/) - Blog Summary: (AI Summaries by Summarizes) The author is not a fan of Apple’s software but thinks their hardware is well-made. A recent rumor might be changing the author’s mind about Apple’s hardware. BeamItDown software, a company that sold contemporary ebooks on iOS devices, had to cease operations because Apple made it impossible for them - [Life Without a Cellphone/Smartphone](https://www.jesse-anderson.com/2011/05/life-without-a-cellphonesmartphone/) - Blog Summary: (AI Summaries by Summarizes) The author recently spent some time without a cell phone after losing their company-provided smartphone. The good side of not having a cell phone was being unreachable and enjoying blissful silence. The author used Google Voice to still receive voicemails and send/receive text messages. The bad side of not - [Changes at Google](https://www.jesse-anderson.com/2011/04/changes-at-google/) - Blog Summary: (AI Summaries by Summarizes) Larry Page, the CEO of Google, is changing the way decisions are made at the company to prevent layers of management from “insulating” engineers and stifling innovation. Google has not released the source code for Android 3.0 (Honeycomb), which has raised concerns about the future of Android’s openness. Google - [Hadoop Book Reviews](https://www.jesse-anderson.com/2011/04/hadoop-book-reviews/) - Blog Summary: (AI Summaries by Summarizes) Two Hadoop books were reviewed: “Hadoop the Definitive Guide 2nd Edition” and “Hadoop In Action” by Tom White and Chuck Lam respectively. “Hadoop the Definitive Guide 2nd Edition” is more focused on programming and has more real-world and applicable code examples. The book goes into better detail about the - [Early Java Stories](https://www.jesse-anderson.com/2011/03/early-java-stories/) - Blog Summary: (AI Summaries by Summarizes) Chuck McManis wrote a post about the early days of Java. The post discusses lesser-known facts about Java’s development. The author notes that once something becomes successful, many people come forward to claim credit for it. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ - [Bold Predictions](https://www.jesse-anderson.com/2011/03/bold-predictions/) - Blog Summary: (AI Summaries by Summarizes) An Android powered Kindle is predicted to make its debut soon. Barnes and Noble’s version of the Android Marketplace lacked the ability for 3rd party developers to create their own apps, creating a void for apps unless the Nook was rooted. Amazon created an Android Marketplace without an official - [Qt for Android](https://www.jesse-anderson.com/2011/02/qt-for-android/) - Blog Summary: (AI Summaries by Summarizes) Qt for Android is in Alpha stage A video tutorial on how to set up, develop, and debug Qt applications on Android is available on the Qt blog The tutorial provides a great start for developers interested in using Qt for Android [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ - [Watson on Jeopardy](https://www.jesse-anderson.com/2011/02/watson-on-jeopardy/) - Blog Summary: (AI Summaries by Summarizes) IBM’s Watson on Jeopardy is an impressive feat that combines several complex domains. Watson’s ability to reduce answers into questions was highlighted when it partially answered a question about a gymnast’s anatomical oddity. Watson’s algorithm likely removes anything other than nouns or names for certain answer types. If the - [Verizon+iPhone=Android Killer?](https://www.jesse-anderson.com/2011/01/verizoniphoneandroid-killer/) - Blog Summary: (AI Summaries by Summarizes) Many people believed that the only reason people were buying Androids was because of AT&T’s poor network. With the release of the Verizon iPhone, buyers can get their iPhone and a good network too. Some pundits believe that this will be the end of Android growth, but the author - [Rooted Nook Color Review](https://www.jesse-anderson.com/2011/01/rooted-nook-color-review/) - Blog Summary: (AI Summaries by Summarizes) The Nook Color (NC) is a locked down device, but rooting it allows full control and access to the Android market. Rooting the NC takes some technical expertise and time, but it is well worth the effort. The NC is just the right size and weight for reading, and - [Kinect On Windows](https://www.jesse-anderson.com/2011/01/kinect-on-windows/) - Blog Summary: (AI Summaries by Summarizes) Installing Kinect on Windows is more difficult than on Linux. The proper order and places to install things are discussed in a Google Groups thread. The Kinect driver can be installed from the SensorKinect project. To use the PrimeSense NITE algorithm, OpenNI, Sensor, and NITE need to be installed - [Hacking the Kinect](https://www.jesse-anderson.com/2010/12/hacking-the-kinect/) - Blog Summary: (AI Summaries by Summarizes) Hacking the Kinect involves connecting it to a computer instead of an Xbox 360 and using open source drivers to control and receive data. The Open Kinect site provides the open source drivers and instructions for controlling and getting data from the depth and color cameras. To turn the - [Pre-Interview (Job Interviews)](https://www.jesse-anderson.com/2010/12/pre-interview-job-interviews/) - Blog Summary: (AI Summaries by Summarizes) HR and Engineering often have a disconnect when it comes to job listings and descriptions. Job listings are often written to be as generic as possible, which can lead to incorrect or impossible requirements. HR may have limited technical knowledge and wants to cast a wide net for applicants. - [The Technical Side of Personal Branding](https://www.jesse-anderson.com/2010/11/the-technical-side-of-personal-branding/) - Blog Summary: (AI Summaries by Summarizes) Personal branding sites can be set up through blogging sites or by setting up your own site with a hosting company. Setting up your own site allows for more customization and the ability to choose your own domain name. To set up your own site, the first step is - [The (Sad) State of Java](https://www.jesse-anderson.com/2010/11/the-sad-state-of-java/) - Blog Summary: (AI Summaries by Summarizes) Oracle’s purchase of Sun Microsystems led to the discontinuation of two high profile open source projects, OpenOffice and OpenSolaris. Oracle is suing Google over the use of Java in Google’s Android operating system, claiming that Java’s license forbids it. Apple has deprecated Java, prohibiting apps written for the new - [The Cost of Bad Code](https://www.jesse-anderson.com/2010/11/the-cost-of-bad-code/) - Blog Summary: (AI Summaries by Summarizes) The 2010 CAST Worldwide Application Software Quality Study shows that poorly written code can be costly. “Good code” or “structural quality” is about how well-architected an application is, not just syntax or non-functional properties. Each line of code has $2.82 of Technical Debt and the average-sized application has $1,000,000 - [Keyboard Productivity](https://www.jesse-anderson.com/2010/11/keyboard-productivity/) - Blog Summary: (AI Summaries by Summarizes) Keep your hands on the keyboard as much as possible to avoid grabbing the mouse repeatedly. Use keyboard shortcuts like Ctrl+X, Ctrl+C, and Ctrl+V for cutting, copying, and pasting text. Hold down the select key to select text and use cursor keys to start selecting the text. Use “Page - [Slightly-more-techinical-than-you Syndrome](https://www.jesse-anderson.com/2010/11/slightly-more-techinical-than-you-syndrome/) - Blog Summary: (AI Summaries by Summarizes) Haystack, a program that claimed to allow uncensored internet access in heavily filtered areas like Iran, was found to be a fraud. The media’s lack of fact-checking allowed the story to spread, potentially putting Iranian users in danger. The inventor of Haystack may have taken advantage of the “Slightly-more-technical-than-you - [iPhone and Android Openness](https://www.jesse-anderson.com/2010/11/iphone-and-android-openness/) - Blog Summary: (AI Summaries by Summarizes) Apple has changed its policies to allow the use of cross-platform compilers like Flash to create iPhone apps. Apple is releasing the App Store Review Guidelines. Android’s openness may have motivated Apple to change their policies. Unity was not that different from Flash, yet apps based on Unity were - [User Interface In Technology](https://www.jesse-anderson.com/2010/11/user-interface-in-technology/) - Blog Summary: (AI Summaries by Summarizes) User interface difficulties have existed for every piece of technology, not just computers. America’s fascination with and use of technology has been ongoing since the American Revolution. Henry Ford’s car salesman had to teach people how to drive their new cars, highlighting the importance of getting feedback from end - [What Happened To Yahoo](https://www.jesse-anderson.com/2010/11/what-happened-to-yahoo/) - Blog Summary: (AI Summaries by Summarizes) Yahoo! was once the search engine of choice before being supplanted by Google. A former Yahoo! employee wrote about why this happened. The quality of programmers at a company is crucial to its success in technology. Once a company’s programmers start to decline in quality, it enters a death - [Code Review (For Interviews)](https://www.jesse-anderson.com/2010/11/code-review-for-interviews/) - Blog Summary: (AI Summaries by Summarizes)It is difficult to have a candidate write a substantial amount of code during an interview.A code sample should be shown to the interviewee and they should be asked to circle the problems and give solutions.The code sample should be between 30-50 lines of code and have various problems with - [Google Fu (For Interviews)](https://www.jesse-anderson.com/2010/11/google-fu-for-interviews/) - Blog Summary: (AI Summaries by Summarizes) Google Fu is the ability to use Google (or another search engine) to find the answer to a question. Google Fu is a very important skill for programmers. Programmers who are unable to properly Google things take a very long time to figure out an issue. Google Fu testing - [Programming Job Interviews](https://www.jesse-anderson.com/2010/11/programming-job-interviews/) - Blog Summary: (AI Summaries by Summarizes) Programming job interviews are often not focused on day-to-day programming tasks or how to program well. Many interview questions are about memorization of answers rather than potential programming skills. The current state of interviews is compared to rejecting a football player because they can’t run fast backwards, which is - [SSH Key Generation Instructions](https://www.jesse-anderson.com/2010/11/ssh-key-generation-instructions/) - Blog Summary: (AI Summaries by Summarizes) SSH keys can be created to allow for SSH or checkouts/commits without password prompts Key generation should be done on the newest possible SSH version to avoid weak keys generated by older versions To create a key on a Linux/Mac box: To create a key on Windows: If there - [Courageous Followership for Programmers](https://www.jesse-anderson.com/2010/11/courageous-followership-for-programmers/) - Blog Summary: (AI Summaries by Summarizes) Programmers’ followership is complex. Poor and technically incompetent management can lead to programmers flouting their boss’ decisions. In some cases, the programmers were technically correct as time showed. Technical decisions made by managers can also be flouted by programmers. All decisions must be adhered to, even if the follower - [Give Chrome a Chance](https://www.jesse-anderson.com/2010/11/20-2/) - Blog Summary: (AI Summaries by Summarizes) The author used to use Firefox but has switched to Chrome and is now hooked. The latest version of Chrome has most of the extensions available on Firefox, including AdBlock. The speed of Chrome is impressive. The author has only used Chrome on Windows. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row - [How to Find Crappy Programmers](https://www.jesse-anderson.com/2010/11/how-to-find-crappy-programmers/) - Blog Summary: (AI Summaries by Summarizes) Someone else has already written a blog post on the topic the author planned to write about. The blog post the author wanted to write about highlights warning signs of euphemisms in job listings. Euphemisms in job listings can be a sign of trouble. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row - [Is The iPad the Next Big Thing?](https://www.jesse-anderson.com/2010/11/is-the-ipad-the-next-big-thing/) - Blog Summary: (AI Summaries by Summarizes) The iPad is fast and responsive, with memory directly on the processor. The touch accuracy is nice, with no problems with on-screen keyboards. Downsides include jabs from Amazon about the Kindle being better in direct sunlight and lighter, as well as the lack of multi-tasking. The iPad feels significantly - [Cross-Platform Is Not Limited to Java](https://www.jesse-anderson.com/2010/11/cross-platform-is-not-limited-to-java/) - Blog Summary: (AI Summaries by Summarizes) Java is a great cross-platform technology, but it may not always be the right choice. QT from Nokia is a good cross-platform GUI framework that is LGPL licensed and feature-rich. WxWidgets is another LGPL GUI framework that is highly recommended by others. Ice from ZeroC is a cross-platform and - [Initiative - The Key to Getting a Job](https://www.jesse-anderson.com/2010/11/initiative-the-key-to-getting-a-job/) - Blog Summary: (AI Summaries by Summarizes) Initiative is key to getting a job, and it should be shown before, during, and after the interview process. Simply knowing someone in the business is not enough to get a job, as one needs to have the necessary skills and knowledge. Taking initiative before the interview can involve - [Great Introduction to Source Control](https://www.jesse-anderson.com/2010/11/great-introduction-to-source-control/) - Blog Summary: (AI Summaries by Summarizes) Joel Spolsky created a tutorial on source control aimed at Mercurial. The tutorial is a great resource for those interested in using distributed version control systems. Most of the concepts covered in the tutorial are applicable to all source control systems. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” - [Analogy for Some Data Structures](https://www.jesse-anderson.com/2010/11/analogy-for-some-data-structures/) - Blog Summary: (AI Summaries by Summarizes) Plumbing troubles served as a good analogy for linked lists and arrays/vectors. Googling for plumbers in the Reno area is like an array where all data is visible at once. None of the plumbers had the part needed, which led to calling another plumber and asking for a referral, - [Rethinking IDE User Interface](https://www.jesse-anderson.com/2010/11/rethinking-ide-user-interface/) - Blog Summary: (AI Summaries by Summarizes) There is a demo of an IDE called “Code Bubbles” that is interesting for debugging and working with well factored code. It may be cumbersome for writing a brand new class and large methods. The IDE is based on Eclipse and only works with Java. There is a beta - [Maybe This Is Why Johnny Can't Code](https://www.jesse-anderson.com/2010/11/maybe-this-is-why-johnny-cant-code/) - Blog Summary: (AI Summaries by Summarizes) A comic was found that explains why Johnny Can’t Code. The comic links to a page that promotes learning to program in 10 years. A software developer was fired from a job after attempting to teach himself C++ in 24 hours using a book. [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row - [Alternatives to MVC](https://www.jesse-anderson.com/2010/11/alternatives-to-mvc/) - Blog Summary: (AI Summaries by Summarizes) The author was asked about alternatives to Model View Controller (MVC) during a presentation. The author found a blog post that discusses the majority of alternatives to MVC. After researching, MVC is still the most commonly used pattern for applications and frameworks. Microsoft is a major exception and uses - [Why Johnny Can't Code](https://www.jesse-anderson.com/2010/11/9-2/) - Blog Summary: (AI Summaries by Summarizes) Finding a decent programmer is difficult due to the number of posers in the software industry. Interviews can be disquieting for both the interviewer and interviewee. Interviewees’ anxiety can give the wrong message or make them prone to mistakes. It is the interviewer’s duty to put the interviewee at - [What Knowledge Gaps Do Self-Taught Programmers Generally Have?](https://www.jesse-anderson.com/2010/11/8-2/) - Blog Summary: (AI Summaries by Summarizes) A recent Ask Slashdot question asked about the knowledge gaps that self-taught programmers generally have. The author of the post was introduced to good programming books early on, including the Gang of Four’s Design Patterns book, which helped them develop good habits. One commenter on the Ask Slashdot question - [Morale the Tangible Intangible](https://www.jesse-anderson.com/2010/11/morale-the-tangible-intangible/) - Blog Summary: (AI Summaries by Summarizes) Morale is both tangible and intangible, and high morale leads to higher productivity. There is no clear indicator of an individual’s morale. Companies often try to raise morale through food and environment. Google provides free food and personal development projects to employees, but the ROI is intangible. Pulling food ## Pages - [Home](https://www.jesse-anderson.com/) - The Big Data and AI Expert for 50% of the Fortune 100. Sharing his years of experience and research to build high-performing data teams. - [Press](https://www.jesse-anderson.com/press/) - Press Jesse has been covered in both mainstream and technical press. This includes prestigious publications such as The Wall Street Journal, BBC, CNN, and NPR. Press Coverage Wall Street Journal, Fox News, The Register, Reno Gazette Journal, The Toronto Star, The Random Fact, Ubuntu CloudBBC, CNN, Slashdot (Again), Gizmodo, Engadget, ArsTecnica, Wired, The Washington Post, The Huffington PostNPR, Radio New Zealand, Australian Broadcasting Company, CBS RadioAOLIn Search of the Data Dream - [Speaking](https://www.jesse-anderson.com/speaking/) - Conference Speaking and Keynoting Jesse has spoken at technology conferences around the world. See how he can bring his highly-rated speaking and presentation style to your conference. Contact me Selected Conference Keynotes Here are some common subjects that we’re giving as keynotes:Pulsar for Kafka PeopleFoundations of Data TeamsWhy Most Data Projects Fail and How to - [Podcast](https://www.jesse-anderson.com/podcast/) - Podcast Jesse is both the host of The Data Dream Team podcast and is a frequent guest on other podcasts. Data Dream Team Podcast Learn about data teams as Jesse hosts industry leaders from around the world The Data Dream Team Podcast brings forward different voices and perspectives to help find a common ground to - [Consulting](https://www.jesse-anderson.com/consulting/) - Consulting and Mentoring Let Jesse bring his extensive experience and research directly to your team. Identify the biggest issues, learn how to resolve them, and get results – guaranteed. Contact Jesse Work With Jesse's Company About Big Data Institute At Big Data Institute, we take a long-term approach to help you build the right team - [About](https://www.jesse-anderson.com/about/) - Who is Jesse? Jesse Anderson is a Data Engineer, Creative Engineer, and Managing Director of Big Data Institute.From nimble startups to established Fortune 100s, Jesse mentors companies globally on Big Data, leveraging cutting-edge tech like Apache Kafka, Hadoop, and Spark.A true innovator in the field, Jesse’s groundbreaking teaching methods have not only earned him recognition - [Elementor #3488](https://www.jesse-anderson.com/elementor-3488/) - Name Name Email Send - [Books](https://www.jesse-anderson.com/books/) - The Thought Leader in Big Data Technology and Team Management Jesse has been published extensively on managing data teams and on deep technical subjects. Contact Jesse A Unified Management Model for Successful Data-Focused Teams Data Teams Learn how to run successful big data projects, how to resource your teams, and how the teams should work - [Blog](https://www.jesse-anderson.com/blog/) - Blog Curious about the latest trends in technology? Jesse’s blog, covering everything from cutting-edge data engineering advancements to innovative organizational design strategies, will keep you informed and inspired. Latest Blog Posts Gemini Batch API for Java November 14, 2025 Read More » Unapologetically Technical Episode 20 – Shane Murray May 13, 2025 Read More » - [Real-Time Systems Sale Page](https://www.jesse-anderson.com/real-time-systems-sale-page/) - What is the next Big Data trend and what should you be doing about it? When people ask me what they should be learning next, I tell them to start learning real-time Big Data systems. Real-time Big Data is something I’ve been focusing on for the past 5+ years. This is because I saw it - [Resources](https://www.jesse-anderson.com/resources/) - Resources Jesse has created extensive resources about big data on technical and organizational topics. Trainings & Certifications Academic Papers Podcast Blog Big Data Engineering Get the skills and training to become a Big Data Engineer ~8 Week Course Extensive and In-Depth 1-on-1 Interactions Data Engineering is growth industry How Managing Data Teams Learn how to - [Search Results](https://www.jesse-anderson.com/search_gcse/) - Search Results - [Contact](https://www.jesse-anderson.com/contact/) - Jesse Anderson +1 775 393 9122 Let's Talk How can I help your data team be successful? I’d love to hear from you. Name Email Address Message Submit - [Trainings & Certifications](https://www.jesse-anderson.com/trainings-certifications/) - Trainings & Certifications Jesse offers live virtual and in-person training and course completion certifications through Big Data Institute. These are for companies seeking to train large groups in their organization. Please note that there are self-guided classes for individuals. Virtual and In-Person Training Big Data Institute Jesse’s virtual and in-person classes are available through - [Privacy Policy](https://www.jesse-anderson.com/privacy-policy/) - Privacy PolicyWho we areOur website address is: https://www.jesse-anderson.com.What personal data we collect and why we collect itContact formsIf you contact us via our contact forms, we will contact you using the email address.Embedded content from other websitesArticles on this site may include embedded content (e.g. videos, images, articles, etc.). Embedded content from other websites behaves - [Academic Papers](https://www.jesse-anderson.com/academic-papers/) - Academic Papers Jesse shares his knowledge, experience, and research in a variety of ways, including academic papers and citations. Academic Papers and Citations Agogo, D. & Anderson, J. (2019). Teaching Tip: The Data Shuffle: Using Playing Cards to Illustrate Data Management Concepts to a Broad Audience. Journal of Information Systems Education, 30(2), 84-96.Kohar, Richard. Basic Discrete Mathematics: Logic, - [Theme Style](https://www.jesse-anderson.com/theme-style/) - Add Your Heading Text Here Add Your Heading Text Here Add Your Heading Text Here Add Your Heading Text Here Add Your Heading Text Here Add Your Heading Text Here Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Pellentesque massa placerat duis ultricies lacus - [Confirming Your Opt-In](https://www.jesse-anderson.com/confirming-your-opt-in/) - [et_pb_section admin_label=”section”] [et_pb_row admin_label=”row”] [et_pb_column type=”4_4″][et_pb_text admin_label=”Text”] Wait! There’s one last thing you need to do. You’re going to receive an email confirming that you want to subscribe to my email list. You will need to click on this activation before you’ll receive anything from me. If you don’t receive this confirmation, you might need - [Switching Careers Confirmation](https://www.jesse-anderson.com/switching-careers-confirmation/) - [et_pb_section fb_built=”1″ _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” max_width=”600px” use_custom_width=”on” custom_width_px=”600px” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text _builder_version=”3.27.4″ header_font=”||||||||” header_line_height=”1.2em” global_colors_info=”{}”] Your link to The Ultimate Guide to Switching Careers to Big Data is on its way [/et_pb_text][/et_pb_column][/et_pb_row][et_pb_row _builder_version=”3.25″ global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_image src=”https://www.jesse-anderson.com/wp-content/uploads/2017/10/cover_400_border.jpg” align=”center” align_tablet=”center” align_phone=”” align_last_edited=”on|desktop” _builder_version=”3.23″ global_colors_info=”{}”][/et_pb_image][/et_pb_column][/et_pb_row][et_pb_row _builder_version=”3.25″ max_width=”800px” use_custom_width=”on” - [MapReduce or Spark Email](https://www.jesse-anderson.com/mapreduce-or-spark-email/) - [et_pb_section fb_built=”1″ admin_label=”section” _builder_version=”3.22″ global_colors_info=”{}”][et_pb_row admin_label=”row” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” max_width=”663px” use_custom_width=”on” custom_width_px=”663px” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text admin_label=”On Its Way” _builder_version=”3.27.4″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” global_colors_info=”{}”] Your free video is on its way! [/et_pb_text][et_pb_image src=”https://www.jesse-anderson.com/wp-content/uploads/2016/03/Screen-Shot-2016-03-14-at-12.06.11-PM.png” align=”center” align_tablet=”center” align_phone=”” align_last_edited=”on|desktop” admin_label=”MR or Spark Image” _builder_version=”3.23″ max_width=”300px” animation_style=”slide” animation_direction=”left” animation_duration=”500ms” animation_intensity_slide=”10%” global_colors_info=”{}”][/et_pb_image][et_pb_text admin_label=”Download” _builder_version=”3.27.4″ background_size=”initial” background_position=”top_left” - [Get Your Data Engineering Dream Job](https://www.jesse-anderson.com/get-your-data-engineering-dream-job/) - [et_pb_section fb_built=”1″ custom_padding_last_edited=”on|desktop” module_id=”header_section” _builder_version=”3.22″ background_color=”#ffffff” custom_margin=”||0px|” custom_padding=”||0px|” custom_padding_tablet=”50px|0|50px|0″ custom_padding_phone=”” transparent_background=”off” global_colors_info=”{}”][et_pb_row padding_mobile=”off” column_padding_mobile=”on” _builder_version=”3.25″ background_size=”initial” background_position=”top_left” background_repeat=”repeat” max_width=”600px” use_custom_width=”on” custom_width_px=”600px” global_colors_info=”{}”][et_pb_column type=”4_4″ _builder_version=”3.25″ custom_padding=”|||” global_colors_info=”{}” custom_padding__hover=”|||”][et_pb_text admin_label=”Get Your Data Engineering Dream Job” _builder_version=”3.27.4″ text_font=”||||||||” text_font_size=”34″ header_font=”||||||||” header_line_height=”1.2em” background_size=”initial” background_position=”top_left” background_repeat=”repeat” text_orientation=”center” module_alignment=”center” header_line_height_last_edited=”off|desktop” global_colors_info=”{}”] Learn Big Data in 8 weeks, make your career switch, - [Should You Learn Spark or MapReduce?](https://www.jesse-anderson.com/gift-spark-or-mapreduce/) - [et_pb_section bb_built=”1″ admin_label=”section” transparent_background=”off” allow_player_pause=”off” inner_shadow=”off” parallax=”off” parallax_method=”off” custom_padding=”|0px||” padding_mobile=”off” make_fullwidth=”off” use_custom_width=”off” width_unit=”on” make_equal=”off” use_custom_gutter=”off”][et_pb_row admin_label=”Row” make_fullwidth=”off” use_custom_width=”off” width_unit=”on” use_custom_gutter=”off” custom_padding=”||0px|” padding_mobile=”off” allow_player_pause=”off” parallax=”off” parallax_method=”off” make_equal=”off” parallax_1=”off” parallax_method_1=”off” parallax_2=”off” parallax_method_2=”off” column_padding_mobile=”on”][et_pb_column type=”3_4″][et_pb_text admin_label=”Gumroad Buy” background_layout=”light” text_orientation=”center” use_border_color=”off” border_color=”#ffffff” border_style=”solid” custom_margin=”|0px||” custom_padding=”|0px||”] [/et_pb_text][/et_pb_column][et_pb_column type=”1_4″][et_pb_blurb admin_label=”Really Free” title=”Really Free” url_new_window=”off” use_icon=”on” font_icon=”%%2%%” icon_color=”#7EBEC5″ use_circle=”off” circle_color=”#7EBEC5″ use_circle_border=”off” circle_border_color=”#7EBEC5″ ## Courses - [Big Data Engineering](https://www.jesse-anderson.com/courses/big-data-engineering/) - Big Data Engineering Get the skills and training to become a Big Data Engineer ~8 Week Course Extensive and In-Depth 1-on-1 Interactions Data Engineering is growth industry How can you get the Big Data skills to do it? I’ve created an 8-week course that takes people like you and makes them Data Engineers. It teaches - [Managing Data Teams](https://www.jesse-anderson.com/courses/managing-data-teams/) - Managing Data Teams Learn how to run and manage effective data teams 16+ Hour Course Real-World Exercises 1-on-1 Interactions Running Data Teams is different than a software engineering or data warehouse team How can you get the knowledge skills to do it right? This is an upcoming course. It will share: How to organize data teams How to - [Real-time Data Engineering](https://www.jesse-anderson.com/courses/real-time-data-pipelines/) - Real-time Data Engineering Take your Big Data skills to the next level with real-time technologies like Apache Spark and Apache Kafka. 16+ Hour Course Real-World Exercises Individual Interactions Real-time Big Data is the Next Industry Trend How can you get the skills to do it? I’ve created a brand-new course for intermediate to advanced Big Data ## Categories - [Uncategorized](https://www.jesse-anderson.com/category/uncategorized/) - [Blog](https://www.jesse-anderson.com/category/blog/) - [Business](https://www.jesse-anderson.com/category/business/) - [Data Engineering](https://www.jesse-anderson.com/category/data-engineering/) - [Featured Talks](https://www.jesse-anderson.com/category/featured-talks/) - [News](https://www.jesse-anderson.com/category/news/) - [Data Engineering is hard](https://www.jesse-anderson.com/category/blog/data-engineering-is-hard/) - Posts about why data engineering is hard and you need to specialize in it. - [Magnum Opus](https://www.jesse-anderson.com/category/blog/magnumopus/) - The blog posts I think represent a great post or project. - [Million Monkeys](https://www.jesse-anderson.com/category/blog/million-monkeys-blog/) - The posts related to the Million Monkeys project where I randomly recreated Shakespeare using Hadoop. - [NNSDG Blog](https://www.jesse-anderson.com/category/blog/nnsdgblog/) - These are the blog posts I make on the Northern Nevada Software Developers Group blog. ## Course Categories - [Learn At Your Own Pace​](https://www.jesse-anderson.com/categories/learn-at-your-own-pace/)