
Many economic factors are at play, Next Generation Europe is going to require strong support from Next Generation Internet in key innovative technologies like AI and Big Data. The capacity to extract the maximum value out of data and create evidence-based policymaking will make the difference as part of the economical recovery (by @jara from HOPU).
This analysis makes a special focus on #DataQuality with new standards as IEEE P2510 led by HOPU, FIWARE solutions as CEF Context Broker for messaging, data exchange; and data marketplaces as the lead by i3-Market https://i3-market.eu/.
To succeed, every modern company will need to be not just a software company but also a data company. There is, of course, some overlap between software and data, but data technologies have their own requirements, tools, and expertise. And some data technologies involve an altogether different approach and mindset – machine learning, for all the discussion about commoditization, is still a very technical area where success often comes in the form of 90-95% prediction accuracy, rather than 100%. This has deep implications for how to build AI products and companies.
In order to foster urban sustainability and economic recovery. We must think about value engineering as we are doing in HOPU with our "Human Oriented approach for Products design". We need to go into urban innovations using the latest techs as AI, IoT and Data-Quality. The decision-makers are going to create new evidence-based policymaking driven by data. NGI with startups as HOPU are enabling a new dimension of tools to guarantee that data is understandable, intuitive and usable. We are here to support sustainable urban development. We really believe our planet sustainability must be supported by all the advances in sensing and AI capabilities.
Examples about data quality and its integration with blockchain are key reference implementations to provide trustability and reliability to AI, via innovative technologies as Blockchain which are offering the trustability medium. Excellent examples from EU are found in Blockchers and Alastria ecosystems, such as visua
Download link of the high definition landscape map.
Let's dig in the landscape technologies and key innovations over the analysis from https://venturebeat.com/2020/10/21/the-2020-data-and-ai-landscape/ and explore how it is translated to EU market.
There’s plenty going on in data infrastructure in 2020. As companies start reaping the benefits of the data/AI initiatives they started over the last few years, they want to do more. They want to process more data, faster and cheaper. They want to deploy more ML models in production. And they want to do more in real-time. Etc.
This raises the bar on data infrastructure (and the teams building/maintaining it) and offers plenty of room for innovation, particularly in a context where the landscape keeps shifting (multi-cloud, etc.).
While those trends are still very much accelerating, here are a few more that are top of mind in 2020:
1. The modern data stack goes mainstream. The concept of “modern data stack” (a set of tools and technologies that enable analytics, particularly for transactional data) has been many years in the making. It started appearing as far back as 2012, with the launch of Redshift, Amazon’s cloud data warehouse.
But over the last couple of years, and perhaps even more so in the last 12 months, the popularity of cloud warehouses has grown explosively, and so has a whole ecosystem of tools and companies around them, going from leading edge to mainstream.
The general idea behind the modern stack is the same as with older technologies: To build a data pipeline you first extract data from a bunch of different sources and store it in a centralized data warehouse before analyzing and visualizing it.
But the big shift has been the enormous scalability and elasticity of cloud data warehouses (Amazon Redshift, Snowflake, Google BigQuery, and Microsoft Synapse, in particular). They have become the cornerstone of the modern, cloud-first data stack and pipeline.
While there are all sorts of data pipelines (more on this later), the industry has been normalizing around a stack that looks something like this, at least for transactional data:
2. ELT starts to replace ELT. Data warehouses used to be expensive and inelastic, so you had to heavily curate the data before loading into the warehouse: first extract data from sources, then transform it into the desired format, and finally load into the warehouse (Extract, Transform, Load or ETL).
In the modern data pipeline, you can extract large amounts of data from multiple data sources and dump it all in the data warehouse without worrying about scale or format, and then transform the data directly inside the data warehouse – in other words, extract, load, and transform (“ELT”).
Key innovations are the introduction in Europe of APACHE NiFi as FIWARE DRACO in FIWARE ecosystem and projects as MIDIH.
3. Data engineering is in the process of getting automated. ETL has traditionally been a highly technical area and largely gave rise to data engineering as a separate discipline. This is still very much the case today with modern tools like Spark that require real technical expertise.
4. Data analysts take a larger role. An interesting consequence of the above is that data analysts are taking on a much more prominent role in data management and analytics.
Data analysts are non-engineers who are proficient in SQL, a language used for managing data held in databases. They may also know some Python, but they are typically not engineers. Sometimes they are a centralized team, sometimes they are embedded in various departments and business units.
Traditionally, data analysts would only handle the last mile of the data pipeline – analytics, business intelligence, and visualization.
Now, because cloud data warehouses are big relational databases (forgive the simplification), data analysts are able to go much deeper into the territory that was traditionally handled by data engineers, leveraging their SQL skills (DBT and others being SQL-based frameworks).
This is good news, as data engineers continue to be rare and expensive. There are many more (10x more?) data analysts, and they are much easier to train.
In addition, there’s a whole wave of new companies building modern, analyst-centric tools to extract insights and intelligence from data in a data warehouse centric paradigm.
For example, there is a new generation of startups building “KPI tools” to extract insights around specific business metrics, or detecting anomalies, including Sisu, Outlier, or Anodot (which started in the observability data world). An example in the smart cities domain and for the urban development, climate change mitigation featured by OECD about how to to create a Clean, Green Disrupting Machines: The role of IoT and AI to improve cities and tackle climate change.
5. Data lakes and data warehouses may be merging. Another trend towards simplification of the data stack is the unification of data lakes and data warehouses. Some (like Databricks) call this trend the “data lakehouse.” Others call it the “Unified Analytics Warehouse.”
Historically, you’ve had data lakes on one side (big repositories for raw data, in a variety of formats, that are low-cost and very scalable but don’t support transactions, data quality, etc.) and then data warehouses on the other side (a lot more structured, with transactional capabilities and more data governance features).
Data lakes have had a lot of use cases for machine learning, whereas data warehouses have supported more transactional analytics and business intelligence.
The net result is that, in many companies, the data stack includes a data lake and sometimes several data warehouses, with many parallel data pipelines.
Companies in the space are now trying to merge the two, with a “best of both worlds” goal and a unified experience for all types of data analytics, including BI and machine learning.
The “Collaborative, Secure, and Replicable Open Source Data Lakes for Smart Cities” (ODALA) is a strategic project to improve data management in cities. European cities and regions from four different countries together with a cluster of private companies and research institutes will leverage open source technologies and digital transformation – for the benefit of public administrations.
ODALA will adopt the European Union Digital Service Infrastructure (DSI), also known as Connecting Europe Facility (CEF) Building Blocks. The building blocks support the creation of a digital single market where cities and companies can connect and share data. This environment is called ‘data lake’ and will allow cities to connect different data sources – static, historical, and real-time data – from diverse departments within the cities.
Connecting city data in a ‘data lake’ will create better insights for city experts and decision-makers. Finally, it will enable them to make better decisions to improve the quality of life of their citizens.
https://oascities.org/odala-developing-the-future-of-smart-cities-communities/
The ODALA project brings together the combined expertise of companies, research, associations and cities. In total, 15 partners from Belgium, Germany, Finland, France, and Spain form ODALA. The partners are:
City of Kiel (Germany, Coordinator), NEC Laboratories Europe (Germany), imec (Belgien), FIWARE Foundation (Germany), City of Heidelberg (Germany), City of Cartagena (Spain), City of Saint-Quentin (France), Faubourg Numérique (France), Flanders Information Agency (Belgium), HOP Ubiquitous (Spain), Digipolis Antwerp (Belgium), Sirus (Belgium), Vero City (Germany), Contrasec (Finland) and Open & Agile Smart Cities (Belgium)
A lot of the trends I’ve mentioned above point toward greater simplicity and approachability of the data stack in the enterprise. However, this move toward simplicity is counterbalanced by an even faster increase in complexity.
The overall volume of data flowing through the enterprise continues to grow an explosive pace. The number of data sources keeps increasing as well, with ever more SaaS tools.
There is not one but many data pipelines operating in parallel in the enterprise. The modern data stack mentioned above is largely focused on the world of transactional data and BI-style analytics. Many machine learning pipelines are altogether different.
There’s also an increasing need for real time streaming technologies, which the modern stack mentioned above is in the very early stages of addressing (it’s very much a batch processing paradigm for now).
For this reason, the more complex tools, including those for micro-batching (Spark) and streaming (Kafka and, increasingly, Pulsar) continue to have a bright future ahead of them. The demand for data engineers who can deploy those technologies at scale is going to continue to increase.
There are several increasingly important categories of tools that are rapidly emerging to handle this complexity and add layers of governance and control to it.
Orchestration engines are seeing a lot of activity. Beyond early entrants like Airflow and Luigi, a second generation of engines has emerged, including Prefect and Dagster, as well as Kedro and Metaflow. Those products are open source workflow management systems, using modern languages (Python) and designed for modern infrastructure that create abstractions to enable automated data processing (scheduling jobs, etc.), and visualize data flows through DAGs (directed acyclic graphs).
Pipeline complexity (as well as other considerations, such as bias mitigation in machine learning) also creates a huge need for DataOps solutions, in particular around data lineage (metadata search and discovery), as highlighted last year, to understand the flow of data and monitor failure points. This is still an emerging area, with so far mostly homegrown (open source) tools built in-house by the big tech leaders: LinkedIn (Datahub), WeWork (Marquez), Lyft (Admunsen), or Uber (Databook). Some promising startups are emerging.
There is a related need for data quality solutions, and we’ve created a new category in this year’s landscape for new companies emerging in the space (see chart).
Overall, data governance continues to be a key requirement for enterprises, whether across the modern data stack mentioned above (ELTG) or machine learning pipelines.
It’s boom time for data science and machine learning platforms (DSML). These platforms are the cornerstone of the deployment of machine learning and AI in the enterprise. The top companies in the space have experienced considerable market traction in the last couple of years and are reaching large scale.
At one end of the spectrum, the big tech companies (GAFAA, Uber, Lyft, LinkedIn etc) continue to show the way. They have become full-fledged AI companies, with AI permeating all their products. This is certainly the case at Facebook (see my conversation with Jerome Pesenti, Head of AI at Facebook). It’s worth nothing that big tech companies contribute a tremendous amount to the AI space, directly through fundamental/applied research and open sourcing, and indirectly as employees leave to start new companies (as a recent example, Tecton.ai was started by the Uber Michelangelo team).
At the other end of the spectrum, there is a large group of non-tech companies that are just starting to dip their toes in earnest into the world of data science, predictive analytics, and ML/AI. Some are just launching their initiatives, while others have been stuck in “AI purgatory” for the last couple of years, as early pilots haven’t been given enough attention or resources to produce meaningful results yet.
Somewhere in the middle, a number of large corporations are starting to see the results of their efforts. They typically embarked years ago on a journey that started with Big Data infrastructure but evolved along the way to include data science and ML/AI.
Those companies are now in the ML/AI deployment phase, reaching a level of maturity where ML/AI gets deployed in production and increasingly embedded into a variety of business applications. The multi-year journey of such companies has looked something like this:
As ML/AI gets deployed in production, several market segments are seeing a lot of activity:
While it will take several more years, ML/AI will ultimately get embedded behind the scenes into most applications, whether provided by a vendor, or built within the enterprise. Your CRM, HR, and ERP software will all have parts running on AI technologies.
Just like Big Data before it, ML/AI, at least in its current form, will disappear as a noteworthy and differentiating concept because it will be everywhere. In other words, it will no longer be spoken of, not because it failed, but because it succeeded.
Source: https://venturebeat.com/2020/10/21/the-2020-data-and-ai-landscape/