]> WP2 – Storing and Querying Very Large Knowledge bases <p>This work package implements the LOD2 knowledge store component needed for managing the Web of Linked Data as a vast database. The starting point is OpenLink Virtuoso and MonetDB on the database side and Sindice on the information retrieval side. The present data volumes handled are around 10 billion RDF triples and the target is over the 1 trillion triples. The approach is scale-out out (physical) complemented with significant scalability improvements in the RDF engine (logical): Physically, when data grows, servers can be added and data redistributed without interruption of service, ibid for server failure. Logically, there is no point answering questions nobody is asking. Therefore the base data is kept as RDF with text indexing and search ranking. Additional inference results or indices for caching joins are made as a by-product of querying; further exploitation of such (partially) materialized inferences through the graph at run-time is to exploit structural correlations in the graphs.</p> 2 D2.1.1 – Initial LOD cloud hosted on the LOD2 Knowledge Store Cluster 2010-11-30 D2.1.2 – State-of-the-Art and LOD2 Laboratory 2011-02-28 D2.1.3 – Intermediate hosted LOD cloud (50B) and knowledge store evaluation 2011-11-30 D2.1.3 - D2.1.5 are performance and deployment targets. The performance and deployment targets are validated with real and synthetic workloads as described in Task 2.1. The baseline is the present LOD cache of about 8 billion triples, expanded with other data as it becomes available. D2.1.4 – Intermediate hosted LOD cloud (250B) and knowledge store evaluation 2012-11-30 D2.1.3 - D2.1.5 are performance and deployment targets. The performance and deployment targets are validated with real and synthetic workloads as described in Task 2.1. The baseline is the present LOD cache of about 8 billion triples, expanded with other data as it becomes available. D2.1.5 – Final hosted LOD cloud (1T) and knowledge store evaluation 2013-11-30 This deliverable is an update of the copy of the Linked Open Data cloud (including all major data sets) hosted on the LOD2 Knowledge Store Cluster as initially deployed in <a href="../../Deliverable/D2.1.1">D2.1.1</>. The deliverable will include benchmark results for major RDF benchmarks, which demonstrate the scalability improvements of the LOD2 Knowledge Store component throughout the project. For this particular deliverable we envision the scalability to reach 1 trillion triples. 2011-08-31 As servers are added or removed, data is migrated to keep a given level of redundancy and load balance. An initial implementation of reusable intermediate join and owl:sameAs inferences is also included in this deliverable. D2.3 – Integration of OpenLink Virtuoso and MonetDB 2011-08-31 D2.4 – Adaptive caching of intermediate joins and inferences 2012-02-29 The results of T2.2 and T2.3 are complete and production strength. The database adapts to workload and configuration changes and creates and discards auxiliary structures as needed. D2.5 – MonetDB release with optimized graph path processing 2012-08-31 The results of T2.5 are involve cutting-edge database research, yet will be available in an open-source release of CWI’s open source database system MonetDB. Where useful, OpenLink will incorporate such features in a later Virtuoso release as well and make them available to the LOD2 Stack via its virtual database feature (effectively embedding MonetDB into Virtuoso). D2.6 – LOD2 knowledge store release with integrated bulk processing features 2012-08-31 This release will add map-reduce style function-shipping-to-cluster-nodes functionality to the knowledge store. D2.7 – LOD2 knowledge store release with enhanced entity ranking 2013-02-28 This release will contain entity ranking functionality to SPARQL queries which going beyond standard ORDER BY clauses. D2.8 – LOD2 knowledge store release with basic geographical indexing 2013-08-31 This release of the LOD2 knowledge store will contain functionality that integrates spatial criteria into search and demonstrates improved scale of spatial queries from adaptive caching. Within this cluster of activities we will develop the technologies necessary to deliver the LOD2 promise – creating an integrated infrastructure for semantic information integration on the Web. In this context, three core challenges need to be addressed: <ol> <li>ensure the knowledge stores scale with the size of the Data Web,</li> <li>provide means for automatic knowledge extraction, alignment, interlinking and</li> <li>enable social collaboration.</li> </ol> M2.1 – LOD2 Knowledge store V1 and hosted LOD cloud 2011-11-30 M2.2 – LOD2 Knowledge store V2 and hosted LOD cloud 2012-11-30 M2.3 – LOD2 Knowledge store V3 and hosted LOD cloud 2013-11-30 WP1 – Requirements, Design and LOD2 Stack Prototype <p>Objectives of this work package are: (1) to develop use case specifications and to collect user requirements by consulting the communities of practice relevant for the LOD2 use cases and additional prospective application scenarios, (2) to identify technical constraints as well as standards, (3) to produce the architecture and the LOD2 Stack design and (4) to produce an early prototype of the LOD2 Stack in the first year.</p> 1 WP3 – Knowledge Base Creation, Enrichment and Repair <p>WP3 contains tasks focused on the transformation of legacy data to RDF and Linked Data and furthermore on the improvement of existing or extracted data especially with respect to schema enrichment and ontology repair. It is complementary to WP4, which is concerned with interlinking several knowledge bases and providing unified views of them. Tasks concerning the triplification of data will be grounded on existing techniques and know-how of the consortium and will be refined during the lifetime of this project and integrated into the LOD2 Stack. Legacy data triplification represents the entry point for legacy systems to participate in the LOD cloud. The members of the Consortium are leading in the development of transformational tools such as Virtuoso Sponger, RDF Views, D2R server, Triplify, and the DBpedia framework, which have received high acceptance in the Linked Data community.</p> 3