Weblog

Berlin SPARQL Benchmark Version 3 and Benchmarking Results

Hi all,

we are happy to announce Version 3 of the BerlinBerlin is the capital city of Germany and is one of the 16 states of Germany. With a population of 3.45 million people, Berlin is Germany's largest city. It is the second most populous city proper and the seventh most populous urban area in the European Union. Located in northeastern Germany, it ... SPARQL Benchmark as well as the results of a benchmark experiment in which we compared the query, load and update performance of Virtuoso, Jena TDB, 4store4store was designed by Steve Harris and developed at Garlik to underpin their Semantic Web applications. It has been providing the base platform for around 3 years. At times holding and running queries over databases of 15GT, supporting a Web application used by thousands of people., BigDataBigdata® is a horizontally scaled storage and computing fabric supporting optional transactions, very high concurrency, and very high aggregate IO rates. Bigdata® was designed from the ground up as a distributed database architecture running over clusters of 100s to 1000s of machines, but can ..., and BigOWLIM using the new benchmark.

The Berlin SPARQL Benchmark Version 3 (BSBM V3) defines three different query mixes that test different capabilities of RDF stores:

  1. The Explore query mix tests the query performance of RDF stores using simple SPARQL 1.0 queries.
  2. The Explore-and-Update query mix test the read and write performance using SPARQL 1.0 SELECT queries as well as SPARQL 1.1 Update queries.
  3. The Business Intelligence query mix consists of complex SPARQL 1.1 queries that rely on aggregation as well as subqueries and each touches large parts of the test dataset.

 The BSBM V3 specification is found at http://www4.wiwiss.fu-berlin.de/bizer/BerlinSPARQLBenchmark/spec/20101129/

 We also conducted a benchmark experiment in which we compared query, load and update performance of Virtuoso, Jena TDB, 4store, BigData, and BigOWLIM using the new benchmark.

 We tested the stores with for 100 million triple and 200 million triple data sets and ran the Explore as well as the Explore-And-Update query mixes.

The results of this experiment are found at  http://www4.wiwiss.fu-berlin.de/bizer/BerlinSPARQLBenchmark/results/V6/index.html

 It is interesting to see that:

  1. Virtuoso dominates the Explore use case for multiple clients.
  2. BigOwlim also shows good multi-client scaling behavior for the 100M dataset.
  3. 4store is the fastest store for the Explore-And-Update query mix.
  4. BigOwlim is able to load the 200m dataset in under 40 minutes, which comes near the bulk load times of relational databases like MySQL.
  5. All stores that we have previously tested with BSBM V2 improved their query performance and load times.

 We also tried to run the Business Intelligence query mix against the stores. BigData and 4store currently do not provide all SPARQL features that are required to run the BI query mix. We thus tried to run the Business Intelligence query mix only against Virtuoso, TDB and BigOwlim. Doing this, we ran into several “technical problems” that prevented us from finishing the tests and from reporting meaningful results. We thus decided to give the store vendors more time to fix and optimize their stores and will run the BI query mix experiment again in about four months (July 2011).

 Thanks a lot to Orri Erling for his proposal to have the Business Intelligence use case and initial queries for the query mix. Lots of thanks also go to Ivan Mikhailov for his in-depth review of the Business Intelligence query mix and for finding several bugs in the queries. We also want to thank Peter Boncz and Hugh Williams for feedback on the new version of the BSBM benchmark.

 We want to thank the store vendors and implementers for helping us to setup and configure their stores for the experiment. Lots of thanks to Andy Seaborne, Ivan Mikhailov, Hugh Williams, Zdravko Tashev, Atanas Kiryakov, Barry Bishop, Bryan Thompson,  Mike Personick and Steve Harris.

More information about the Berlin SPARQL Benchmark is found at http://www4.wiwiss.fu-berlin.de/bizer/BerlinSPARQLBenchmark/

Posted in WP2 – Storing and Querying

One Response to Berlin SPARQL Benchmark Version 3 and Benchmarking Results

  1. Pingback: BSBM Reduced Query Mix – 36,608 QMpH » bigdata®

Leave a Reply

Your email address will not be published. Required fields are marked *

*

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <strike> <strong>


join our monthly webinars

Public Mailinglist & Newsletter

Please subscribe me to the LOD2 mailinglist.
my email address
my name (optional)
goto archive

Follow


Follow lod2project on Twitter

RSS Twitter

  • linkeddata
    始まりまどか☆ RT @mhatta: RT @IMAGEDRIVE: #lodjp 第6回LinkedData勉強会「オープンライセンス」19:00〜 Ust → http://t.co/2kJZMkxvzC […]
  • linkeddata
    RT @higa4: 本日の資料です。 「ODCライセンスのOpenStreetMapでの導入事例紹介」 http://t.co/MvEPDK4jXU #LOD #LinkedData #LODjp #OKFj #OSMFJ […]
  • linkeddata
    これからこれ見ます! #lodjp で資料が公開されている模様. RT @synobu: 第6回 LinkedData勉強会「オープンライセンスについて」のUstream配信を行っています. #lodjp http://t.co/qBABOjxpUV […]