The Open Call 4 NGI Searchers are announced


May 13 2024

OC4_outcome.png
Date: Monday 13 May 2024
Time: 09:45-12:00
Meeting: OC4 KickOff Online Conference

The fourth Open Call for NGI Search has received 227 applications from 29 countries!

The following NGI Search project beneficiaries have been introduced. Listen to their podcasts, using the reader below: 

AI-Generated Code Search

/get/bookletCode/attachment?file=Sticker_matchcode.png

AI-Generated Code Search enhances trust for users searching for code on the internet, especially for both AI-generated code found on the web and code created by a generative AI engine or LLM. This tool provides information and validation on the derivative open source origin of AI-generated code to identify its vulnerabilities and licenses to mitigate any software supply chain integrity or security risks associated with using AI-generated code.

Listen to the podcast presentation:

00:00 00:00

Carbon.txt

/get/bookletCode/attachment?file=StickerGreenWebFoundation.png

Carbon.txt proposes a single place to look on any domain - /carbon.txt - for public machine-readable sustainability data relating to that company. This standardised, distributed approach uses existing internet governance structures and industry norms to efficiently surface data that companies have to publish according to national laws. We envisage a data ecosystem growing around this that will support action on the defining crisis of our time - the climate crisis.

Listen to the podcast presentation:

00:00 00:00

COGNITREK (Driving Cognitive Accessibility) 

/get/bookletCode/attachment?file=StickerCognitrek_V2.png

Cognitrek is an AI-powered tool that simplifies complex documents and adapts them to students' needs. It extracts and transforms information to make knowledge more accessible, including features that support individuals with learning difficulties. The platform ensures transparency, traceability, and complies with EU ethical and legal standards for AI, including data protection regulations.

Listen to the podcast presentation:

00:00 00:00

COPS

StickerEmpty.png

COPS addresses the problem of private search by combining the functionality of the OpenSearch ( https://opensearch.org/ ) search and analytics framework and the security primitives of confidential computing. The innovation behind this will enable data owners to maintain verifiable control over their data and ensure the protection of data in use.

Listen to the podcast presentation:

00:00 00:00

Data space search engine (EDC) 

/get/bookletCode/attachment?file=StickerStartinBlox.png

Today, companies are only able to share data with each other on a one-to-one basis. To share data, they have to copy-paste it using the Eclipse Data Components (EDC) Connector.

We're extending the EDC Connector to enable indexing, search and discovery of decentralized data among B2B partners. This will bring down the cost of sharing data to a couple of minutes per dataset, will enhance sovereignty over the shared data and will enable B2B data discovery.

Listen to the podcast presentation:

00:00 00:00

Datami

/get/bookletCode/attachment?file=Sticker_Datami.png

Datami PRO empowers small organizations to publish datasets as shareable visualizations that communities can contribute to. Built on the free Datami project, it will be a no-code, open-source platform allowing users to structure datasets into data packages, customize web components (maps, charts, cards, filters), manage contributions, and publish to open data catalogs. Datami PRO’s affordability, flexibility, and respect of standards will make it the simplest tool for anyone's open data project.

Listen to the podcast presentation:

00:00 00:00

Ensuring Fairness in Democratic AI

/get/bookletCode/attachment?file=Sticker_Ensuring-Fairness.png

Matt Stempeck and Eticas are working together to help the builders of civic and democratic AI systems use our data to test whether their systems are in line with best practices regarding fairness. Civic AI builders must prove their democratic principles in their actual technical specifications, not just their intentions. We will help them do just that, and release open data and open libraries, tools, and platforms so others can test their systems for operational fairness.

Listen to the podcast presentation:

thubnail.png
Video presentation - 15 January 2025

Fediverse Discovery Providers

StickerEmpty.png

This project explores the possibilities for better search and discovery on the Fediverse in the form of an optional, pluggable service. This service should be decentralized, independent of any one specific Fediverse service and respect user choice and privacy.

Listen to the podcast presentation:

00:00 00:00

MOUSSE (Metadata fOcUsed Semantic Search Engine) 

/get/bookletCode/attachment?file=Sticker_Mousse.png

Metadata fOcUsed Semantic Search Engine is a semantic search tool powered by advanced large language models (LLMs). It is designed to enhance data discovery by focusing on metadata, making it particularly suitable for large, diverse datasets that lack structured ontologies.

Listen to the podcast presentation:

00:00 00:00

Neural Datafari

/get/bookletCode/attachment?file=StickerNeuralDatafari.png

Adding RAG and vector search to Datafari for a neural end to end search solution, providing heterogeneous crawling capabilities up to a search user interface, empowering citizens and organisations with a privacy preserving complete search solution. Powered by Apache ManifoldCF and Apache Solr.

Listen to the podcast presentation:

00:00 00:00

On My Disk: Web UI

/get/bookletCode/attachment?file=StickerOnMyDisk.png

In this second NGI Search project, the team main milestone has been the development of a user-friendly website hosting feature that allows non-technical users to easily host static sites from their personal cloud. The combination of the On My Disk private cloud solution with the PeARS search engine offers each individual a way to store their content on their private devices, and to make it available to others searching information across the entire network. These efforts should foster more individuals and organisations to design the search engine they need for their community and run it themselves.

Listen to the podcast presentation:

00:00 00:00

Open Data Deep Search

StickerEmpty.png

ODDS (Open Data Deep Search) is an AI based tool that acts as a data analyst: 

  • it finds the correct and relevant datasets 
  • understands the data format and structure 
  • queries the data to find the exact information requested 
  • and provides clear answers to the user’s questions.

Listen to the podcast presentation:

00:00 00:00

Trallie

/get/bookletCode/attachment?file=Sticker_Trallie.png

Trallie is a framework that leverages the power of large language models (LLMs) to reimagine information extraction (IE) with or without user guidelines. Instead of relying on labelled examples or copious amounts of training on your data collection, Trallie requires very few or no representative examples of the data to convert into a structured format. Trallie operates with three core objectives:

  • Minimal manual annotation 
  • Flexibility with document formats and languages 
  • Accessibility to non-expert users.

Listen to the podcast presentation:

00:00 00:00

EU programme:  HORIZON-CL4-2021-HUMAN-01  

flagEU.svg

This project has received funding from the European Union’s Horizon Europe research and innovation programme under the grant agreement 101069364 and it is framed under Next Generation Internet Initiative.

Follow us on social media

linkedin.svgNGISearch_Logo_Icon-circle-N-rgb.svgmastodon-circle.svg

XWiki Enterprise 16.10.17 - Documentation