Cloudera公司已經推出了基於Hadoop平臺的查詢統計分析工具Impala,只要熟悉SQL,就能夠熟練地使用Impala來執行查詢與分析的功能。不過Impala的SQL和關係數據庫的SQL仍是有一點微妙地不一樣的。
下面,咱們設計一個表,經過該表中的數據,來將SQL查詢與統計的語句,使用Solr查詢的方式來與SQL查詢對應。這個翻譯的過程,是很是有趣的,你能夠看到Solr一些很不錯的功能。
用來示例的表結構設計,如圖所示:
下面,咱們經過給出一些SQL查詢統計語句,而後對應翻譯成Solr查詢語句,而後對比結果
查詢對比條件組合查詢SQL查詢語句:html
SELECT log_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_type
數據庫
FROM v_i_event
函數
WHERE prov_id = 1 AND net_type = 1 AND area_id = 10304 AND time_type = 1 AND time_id >= 20130801 AND time_id <= 20130815
工具
ORDER BY log_id LIMIT 10;oop
查詢結果,如圖所示:
Solr查詢URL:spa
http://slave1:8888/solr-cloud/i_event/select?q=*:*&fl=log_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_type&fq=prov_id:1 AND net_type:1 AND area_id:10304 AND time_type:1 AND time_id:[20130801 TO 20130815]&sort=log_id asc&start=0&rows=10翻譯
查詢結果,以下所示:設計
<response>
code
<lst name="responseHeader">
htm
<int name="status">0</int>
<int name="QTime">4</int>
</lst>
<result name="response" numFound="77" start="0">
<doc>
<int name="log_id">6827</int>
<long name="start_time">1375072117</long>
<long name="end_time">1375081683</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10304</int>
<int name="idt_id">11002</int>
<int name="cnt">0</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6827</int>
<long name="start_time">1375072117</long>
<long name="end_time">1375081683</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10304</int>
<int name="idt_id">11000</int>
<int name="cnt">0</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6851</int>
<long name="start_time">1375142158</long>
<long name="end_time">1375146391</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10304</int>
<int name="idt_id">14001</int>
<int name="cnt">5</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6851</int>
<long name="start_time">1375142158</long>
<long name="end_time">1375146391</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10304</int>
<int name="idt_id">11002</int>
<int name="cnt">23</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6851</int>
<long name="start_time">1375142158</long>
<long name="end_time">1375146391</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10304</int>
<int name="idt_id">10200</int>
<int name="cnt">55</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6851</int>
<long name="start_time">1375142158</long>
<long name="end_time">1375146391</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10304</int>
<int name="idt_id">14000</int>
<int name="cnt">4</int>
對比上面結果,除了根據idt_id排序方式不一樣之外(Impala是升序,Solr是降序),其餘是相同的。
單個字段分組統計
SQL查詢語句:
SELECT prov_id, SUM(cnt) AS sum_cnt, AVG(cnt) AS avg_cnt, MAX(cnt) AS max_cnt, MIN(cnt) AS min_cnt, COUNT(cnt) AS count_cnt
FROM v_i_event
GROUP BY prov_id;
查詢結果,如圖所示:
Solr查詢URL:
http://slave1:8888/solr-cloud/i_event/select?q=*:*&stats=true&stats.field=cnt&rows=0&indent=true&stats.facet=prov_id
查詢結果,以下所示:
<response>
<lst name="responseHeader">
<int name="status">0</int>
<int name="QTime">2</int>
</lst>
<result name="response" numFound="4088" start="0"></result>
<lst name="stats">
<lst name="stats_fields">
<lst name="cnt">
<double name="min">0.0</double>
<double name="max">1258.0</double>
<long name="count">4088</long>
<long name="missing">0</long>
<double name="sum">32587.0</double>
<double name="sumOfSquares">9170559.0</double>
<double name="mean">7.971379647749511</double>
<double name="stddev">46.69344567709268</double>
<lst name="facets" />
</lst>
</lst>
</lst>
</response>
對比查詢結果,Solr提供了更多的統計項,如標準差(stddev)等,與SQL查詢結果是一致的。
IN條件查詢SQL查詢語句:
[cde]SELECT log_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_typFROM v_i_eventWHERE prov_id = 1 AND net_type = 1 ANDcity_id IN(106,103) AND idt_id IN(12011,5004,6051,6056,8002) AND time_type = 1AND time_id >= 20130801 AND time_id <= 20130815ORDER BY log_id, start_time DESC LIMIT 10;
[/code]查詢結果,如圖所示:
Solr查詢URL:
http://slave1:8888/solr-cloud/i_event/select?q=*:*&fl=log_id,start_time,end_time,prov_id,city_id,area_id,idt_id, cnt,net_type&fq=prov_id:1 AND net_type:1 AND (city_id:106 OR city_id:103) AND (idt_id:12011 OR idt_id:5004 OR idt_id:6051 OR idt_id:6056 OR idt_id:8002) AND time_type:1 AND time_id:[20130801 TO 20130815]&sort=log_id asc ,start_time desc&start=0&rows=10
或
http://slave1:8888/solr-cloud/i_event/select?q=*:*&fl=log_id,start_time,end_time,prov_id,city_id,area_id,idt_id, cnt ,net_type&fq=prov_id:1&fq=net_type:1&fq=(city_id:106 OR city_id:103)&fq=(idt_id:12011 OR idt_id:5004 OR idt_id:6051 OR idt_id:6056 OR idt_id:8002)&fq=time_type:1&fq=time_id:[20130801 TO 20130815]&sort=log_id asc,start_time desc&start=0&rows=10
查詢結果,以下所示:
<response>
<lst name="responseHeader">
<int name="status">0</int>
<int name="QTime">6</int>
</lst>
<result name="response" numFound="63" start="0">
<doc>
<int name="log_id">6553</int>
<long name="start_time">1374054184</long>
<long name="end_time">1374054254</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10307</int>
<int name="idt_id">12011</int>
<int name="cnt">0</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6553</int>
<long name="start_time">1374054184</long>
<long name="end_time">1374054254</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10307</int>
<int name="idt_id">5004</int>
<int name="cnt">2</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6555</int>
<long name="start_time">1374055060</long>
<long name="end_time">1374055158</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">70104</int>
<int name="idt_id">5004</int>
<int name="cnt">3</int>
<int name="net_type">1</int>
對比查詢結果,是一致的。
開區間範圍條件查詢SQL查詢語句:
SELECTlog_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_typeFROM v_i_eventWHERE net_type = 1 AND idt_idIN(12011,5004,6051,6056,8002) AND time_type = 1 AND start_time >= 1373598465AND end_time < 1374055254
ORDER BY log_id, start_time, idt_id DESCLIMIT 30;查詢結果,如圖所示:
Solr查詢URL:
http://slave1:8888/solr-cloud/i_event/select?q=*:*&fl=log_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_type&fq=net_type:1 AND (idt_id:12011 OR idt_id:5004 OR idt_id:6051 OR idt_id:6056 OR idt_id:8002) AND time_type:1 AND start_time:[1373598465 TO 1374055254]&fq =-start_time:1374055254&sort=log_id asc,start_time asc,idt_id desc&start=0&rows=30
或
http://slave1:8888/solr-cloud/i_event/select?q=*:*&fl=log_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_type&fq=net_type:1 AND (idt_id:12011 OR idt_id:5004 OR idt_id:6051 OR idt_id:6056 OR idt_id:8002) AND time_type:1 AND start_time:[1373598465 TO 1374055254] AND -start_time:1374055254&sort=log_id asc,start_time asc,idt_id desc&start=0&rows=30
或
http://slave1:8888/solr-cloud/i_event/select?q=*:*&fl=log_id,start_time,end_time,prov_id,city_id,area_id,idt_id,cnt,net_type&fq=net_type:1&fq=idt_id:12011 OR idt_id:5004 OR idt_id:6051 OR idt_id:6056 OR idt_id:8002&fq =time_type:1&fq=start_time:[1373598465 TO 1374055254]&fq =-start_time:1374055254&sort=log_id asc,start_time asc,idt_id desc&start=0&rows=30
查詢結果,以下所示:
<response>
<lst name="responseHeader">
<int name="status">0</int>
<int name="QTime">5</int>
</lst>
<result name="response" numFound="4" start="0">
<doc>
<int name="log_id">6553</int>
<long name="start_time">1374054184</long>
<long name="end_time">1374054254</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10307</int>
<int name="idt_id">12011</int>
<int name="cnt">0</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6553</int>
<long name="start_time">1374054184</long>
<long name="end_time">1374054254</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">10307</int>
<int name="cnt">2</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6555</int>
<long name="start_time">1374055060</long>
<long name="end_time">1374055158</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">70104</int>
<int name="idt_id">12011</int>
<int name="cnt">0</int>
<int name="net_type">1</int>
</doc>
<doc>
<int name="log_id">6555</int>
<long name="start_time">1374055060</long>
<long name="end_time">1374055158</long>
<int name="prov_id">1</int>
<int name="city_id">103</int>
<int name="area_id">70104</int>
<int name="idt_id">5004</int>
<int name="cnt">3</int>
<int name="net_type">1</int>
</doc>
</result>
</response>
多個字段分組統計(只支持count函數)SQL查詢語句:SELECT city_id, area_id, COUNT(cnt) AScount_cntFROM v_i_eventWHERE prov_id = 1 AND net_type = 1GROUP BY city_id, area_id;查詢結果,如圖所示:
Solr查詢URL:
http://slave1:8888/solr-cloud/i_event/select?q=*:*&facet=true&facet.pivot=city_id,area_id&fq=prov_id:1 AND net_type:1&rows=0&indent=true
對比上面結果,Solr查詢結果,須要從上面的各組中進行合併,獲得最終的統計結果,結果和SQL結果是一致的。
多個字段分組統計(支持count、sum、max、min等函數)一次對多個字段進行獨立分組統計,Solr能夠很好的支持。這至關於執行兩個帶有GROUP BY子句的SQL,這兩個GROUP BY分別只對一個字段進行彙總統計。
SQL查詢語句:
SELECT city_id, area_id, COUNT(cnt) AS count_cnt
FROM v_i_event
WHERE prov_id = 1 AND net_type = 1
GROUP BY city_id;
SELECT city_id, area_id, COUNT(cnt) AS count_cnt
FROM v_i_event
WHERE prov_id = 1 AND net_type = 1
GROUP BY area_id;
複製代碼
查詢結果,再也不顯示。
Solr查詢URL:
>http://slave1:8888/solr-cloud/i_event/select?q=*:*&stats=true&stats.field=cnt&f.cnt.stats.facet=city_id&&f.cnt.stats.facet=area_id&fq=prov_id:1 AND net_type:1&rows=0&indent=true
查詢結果,以下所示:
<response>
<lst name="responseHeader">
<int name="status">0</int>
<int name="QTime">72</int>
</lst>
<result name="response" numFound="1171" start="0"></result>
<lst name="facet_counts">
<lst name="facet_queries" />
<lst name="facet_fields" />
<lst name="facet_dates" />
<lst name="facet_ranges" />
<lst name="facet_pivot">
<arr name="city_id,area_id">
<lst>
<str name="field">city_id</str>
<int name="value">103</int>
<int name="count">678</int>
<arr name="pivot">
<lst>
<str name="field">area_id</str>
<int name="value">10307</int>
<int name="count">298</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10315</int>
<int name="count">120</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10317</int>
<int name="count">86</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10304</int>
<int name="count">67</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10310</int>
<int name="count">49</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">70104</int>
<int name="count">48</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10308</int>
<int name="count">6</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">0</int>
<int name="count">2</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10311</int>
<int name="count">2</int>
</lst>
</arr>
</lst>
<lst>
<str name="field">city_id</str>
<int name="value">0</int>
<int name="count">463</int>
<arr name="pivot">
<lst>
<str name="field">area_id</str>
<int name="value">0</int>
<int name="count">395</int>
</lst>
<lst>
<str name="field">area_id</str>
<int name="value">10307</int>
<int name="count">68</int>
複製代碼
對比上面結果,Solr查詢結果,須要從上面的各組中進行合併,獲得最終的統計結果,結果和SQL結果是一致的。
多個字段聯合分組統計(支持count、sum、max、min等函數)SQL查詢語句:SELECT city_id, area_id, SUM(cnt) ASsum_cnt, AVG(cnt) AS avg_cnt, MAX(cnt) AS max_cnt, MIN(cnt) AS min_cnt,COUNT(cnt) AS count_cntFROM v_i_eventWHERE prov_id = 1 AND net_type = 1GROUP BY city_id, area_id;
查詢結果,如圖所示:
Solr目前不能簡單的支持這種查詢,若是想要知足這種查詢統計,須要在schema的設計上,將一個字段設置爲多值,而後經過多個值進行分組統計。若是應用中查詢統計分析的模式比較固定,
預先知道哪些字段會用於聯合分組統計,徹底能夠在設計的時候,考慮設置多值字段來知足這種需求。
感興趣的讀者,還能夠看看這裏:基於Solr DIH實現MySQL表數據全量索引和增量索引