主頁 > 資料庫 > Elasticsearch 常用的聚合操作

Elasticsearch 常用的聚合操作

2021-06-09 20:00:14 資料庫

Aggregation 簡介

ps : 本篇文章 Elasticsearch 和 Kibana 版本為 7.10.1,如果版本不一致請查看官方檔案,避免誤導!

聚合框架有助于基于搜索查詢提供聚合資料,它基于稱為聚合的簡單構建塊,可以組合以構建復雜的資料摘要,

Elasticsearch 將聚合分為三類:

  • Metric (指標聚合)

    從欄位值計算度量的聚合,例如最大、最小、總和和平均值,

  • Bucket (桶聚合)

    根據欄位值、范圍或其他條件將檔案分組為桶(也稱為箱),類似于關系型資料庫中的group by,

  • Pipeline (管道聚合)

    從其他聚合而不是檔案或欄位中獲取輸入的聚合,

聚合可以將我們的資料匯總為指標,統計或其他分析資訊,使用聚合可以為我們帶來的好處:

  • 我的網站的平均加載時間是多少?
  • 根據交易量,誰是我最有價值的客戶?
  • 什么會被視為我網路上的大檔案?
  • 每個產品類別中有多少個產品?

資料準備

創建索引

DELETE twitter

PUT twitter
{
    "settings": {
        "number_of_shards": 2,
        "number_of_replicas": 1
    }, 
    "mappings": {
        "properties": {
            "birthday": {
                "type": "date"
            },
            "address": {
                "type": "text",
                "fields": {
                    "keyword": {
                        "type": "keyword",
                        "ignore_above": 256
                    }
                }
            },
            "age": {
                "type": "long"
            },
            "city": {
                "type": "keyword"
            },
            "country": {
                "type": "keyword"
            },
            "location": {
                "type": "geo_point"
            },
            "message": {
                "type": "text",
                "fields": {
                    "keyword": {
                        "type": "keyword",
                        "ignore_above": 256
                    }
                }
            },
            "province": {
                "type": "keyword"
            },
            "uid": {
                "type": "long"
            },
            "user": {
                "type": "text",
                "fields": {
                    "keyword": {
                        "type": "keyword",
                        "ignore_above": 256
                    }
                }
            }
        }
    }
}

匯入資料

使用 Bulk API 將資料匯入到 Elasticsearch 中:

POST _bulk
{"index":{"_index":"twitter","_id":1}}
{"user":"張三","message":"今兒天氣不錯啊,出去轉轉去","uid":2,"age":20,"city":"北京","province":"北京","country":"中國","address":"中國北京市海淀區","location":{"lat":"39.970718","lon":"116.325747"}, "birthday": "1999-04-01"}
{"index":{"_index":"twitter","_id":2}}
{"user":"老劉","message":"出發,下一站云南!","uid":3,"age":22,"city":"北京","province":"北京","country":"中國","address":"中國北京市東城區臺基廠三條3號","location":{"lat":"39.904313","lon":"116.412754"}, "birthday": "1997-04-01"}
{"index":{"_index":"twitter","_id":3}}
{"user":"李四","message":"happy birthday!","uid":4,"age":25,"city":"北京","province":"北京","country":"中國","address":"中國北京市東城區","location":{"lat":"39.893801","lon":"116.408986"}, "birthday": "1994-04-01"}
{"index":{"_index":"twitter","_id":4}}
{"user":"老賈","message":"123,gogogo","uid":5,"age":30,"city":"北京","province":"北京","country":"中國","address":"中國北京市朝陽區建國門","location":{"lat":"39.718256","lon":"116.367910"}, "birthday": "1989-04-01"}
{"index":{"_index":"twitter","_id":5}}
{"user":"老王","message":"Happy BirthDay My Friend!","uid":6,"age":26,"city":"北京","province":"北京","country":"中國","address":"中國北京市朝陽區國貿","location":{"lat":"39.918256","lon":"116.467910"}, "birthday": "1993-04-01"}
{"index":{"_index":"twitter","_id":6}}
{"user":"老吳","message":"好友來了都今天我生日,好友來了,什么 birthday happy 就成!","uid":7,"age":28,"city":"上海","province":"上海","country":"中國","address":"中國上海市閔行區","location":{"lat":"31.175927","lon":"121.383328"}, "birthday": "1991-04-01"}

注意:并不是所有欄位都可以用來做聚合,一般來說,只有具有 keyword或者數值型別的欄位是可以用來做聚合,

我們可以通過 _field_cat 命令還查詢檔案中的欄位是否可以作為聚合:

GET twitter/_field_caps?fields=message,age,province,city.keyword

從結果我們可以看到四個欄位都可以用來做搜索的,但是只有 agecity.keyword才可以用來做聚合

{
  "indices" : [
    "twitter"
  ],
  "fields" : {
    "province" : {
      "text" : {
        "type" : "text",
        "searchable" : true,
        "aggregatable" : false
      }
    },
    "message" : {
      "text" : {
        "type" : "text",
        "searchable" : true,
        "aggregatable" : false
      }
    },
    "city.keyword" : {
      "keyword" : {
        "type" : "keyword",
        "searchable" : true,
        "aggregatable" : true
      }
    },
    "age" : {
      "long" : {
        "type" : "long",
        "searchable" : true,
        "aggregatable" : true
      }
    }
  }
}

searchable

是否為所有索引上的搜索都索引了該欄位,

aggregatable

是否可以在所有索引上匯總此欄位,

indices

該欄位具有相同型別族的索引串列;如果所有索引具有相同的型別族,則為null,

non_searchable_indices

該欄位不可搜索的索引串列;如果所有索引對該欄位的定義相同,則為null,

non_aggregatable_indices

該欄位不可聚合的索引串列;如果所有索引對該欄位的定義相同,則為null,

聚合操作 語法

"aggregations" : {
    "<aggregation_name>" : { <!--聚合的名字 -->
        "<aggregation_type>" : { <!--聚合的型別 -->
            <aggregation_body> <!--聚合體:對哪些欄位進行聚合 -->
        }
        [,"meta" : {  [<meta_data_body>] } ]? <!--元 -->
        [,"aggregations" : { [<sub_aggregation>]+ } ]? <!--在聚合里面在定義子聚合 -->
    }
    [,"<aggregation_name_2>" : { ... } ]*<!--聚合的名字 -->
}

上面的 aggregation 可以使用 aggs 來代替

Metric 聚合操作

Avg Sum Max Min 聚合

Avg Aggregation : 一個單值度量聚合,計算從聚合檔案中提取的數值的平均值,

Sum Aggregation :sum聚合對從聚合檔案中提取的數值進行匯總的單值度量,

Max Aggregation :一個單值度量聚合,用于跟蹤并回傳從聚合檔案中提取的數值中的最大值,

Min Aggregation :一個單值度量聚合,用于跟蹤并回傳從聚合檔案中提取的數值中的最小值,

這些值可以從檔案中的特定數字欄位中提取,也可以由提供的腳本生成,

查詢 twitter 索引下檔案 age 的 平均值、總和、最大值及最小值:

GET twitter/_search?size=0
{
  "aggs": {
    "age_avg": {
      "avg": {
        "field": "age"
      }
    },
    "age_sum":{
      "sum": {
        "field": "age"
      }
    },
    "age_max":{
      "max": {
        "field": "age"
      }
    },
    "age_min":{
      "min": {
        "field": "age"
      }
    }
  }
}

回傳結果:

{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 6,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : [ ]
  },
  "aggregations" : {
    "age_sum" : {
      "value" : 151.0
    },
    "age_min" : {
      "value" : 20.0
    },
    "age_avg" : {
      "value" : 25.166666666666668
    },
    "age_max" : {
      "value" : 30.0
    }
  }
}

Stats 聚合

資料聚合一個多值指標聚合,它根據從聚合檔案中提取的數值計算統計資訊,

回傳的統計資料包括:最小值,最大值,和;

匯總所有檔案的年齡統計

GET twitter/_search?size=0
{
  "query": {
    "match": {
      "city": "北京"
    }
  }, 
  "aggs": {
    "age_stats": {
      "stats": {
        "field": "age"
      }
    }
  }
}

回傳結果

{
    "took" : 0,
    "timed_out" : false,
    "_shards" : {
        "total" : 1,
        "successful" : 1,
        "skipped" : 0,
        "failed" : 0
    },
    "hits" : {
        "total" : {
            "value" : 5,
            "relation" : "eq"
        },
        "max_score" : null,
        "hits" : [ ]
    },
    "aggregations" : {
        "age_stats" : {
            "count" : 5,
            "min" : 20.0,
            "max" : 30.0,
            "avg" : 24.6,
            "sum" : 123.0
        }
    }
}

Bucket 聚合操作

Range 聚合(multi-bucket)

基于多桶值源的聚合,可以定義一組范圍(每個范圍代表一個桶),在聚合程序中,將從每個存盤區范圍中檢查并從檔案中提取值

注意:此聚合包含每個范圍的 from 值,但不包括 to 值,

將年齡進行分段,查詢不同年齡段的用戶:

GET twitter/_search
{
    "size": 0, 
    "aggs": {
        "age_range": {
            "range": {
                "field": "age",
                "ranges": [
                    {
                        "from": 20,
                        "to": 22
                    },
                    {
                        "from": 22,
                        "to": 25
                    },
                    {
                        "from": 25,
                        "to": 30
                    }
                ]
            }
        }
    }
}

上面我們使用 range 型別的聚合,定義了不同的年齡段,通過上面的查詢,得到了不同年齡段的 bucket,并且因為是針對聚合,我們并不關心回傳的結果,通過 size=0 忽略了回傳結果,得到了以下輸出:

{
    "took" : 2,
    "timed_out" : false,
    "_shards" : {
        "total" : 1,
        "successful" : 1,
        "skipped" : 0,
        "failed" : 0
    },
    "hits" : {
        "total" : {
            "value" : 6,
            "relation" : "eq"
        },
        "max_score" : null,
        "hits" : [ ]
    },
    "aggregations" : {
        "age_range" : {
            "buckets" : [
                {
                    "key" : "20.0-22.0",
                    "from" : 20.0,
                    "to" : 22.0,
                    "doc_count" : 1
                },
                {
                    "key" : "22.0-25.0",
                    "from" : 22.0,
                    "to" : 25.0,
                    "doc_count" : 1
                },
                {
                    "key" : "25.0-30.0",
                    "from" : 25.0,
                    "to" : 30.0,
                    "doc_count" : 3
                }
            ]
        }
    }
}

Sub-aggregation

在聚合的內部嵌套一個聚合,

在 range 操作之中,我們可以做 sub-aggregation,分別來計算它們的平均年齡、最大以及最小的年齡!

GET twitter/_search
{
    "size": 0, 
    "aggs": {
        "age_range": {
            "range": {
                "field": "age",
                "ranges": [
                    {
                        "from": 20,
                        "to": 22
                    },
                    {
                        "from": 22,
                        "to": 25
                    },
                    {
                        "from": 25,
                        "to": 30
                    }
                ]
            },
            "aggs": {
                "age_avg": {
                    "avg": {
                        "field": "age"
                    }
                },
                "age_min":{
                    "min": {
                        "field": "age"
                    }
                },
                "age_max":{
                    "max": {
                        "field": "age"
                    }
                }
            }
        }
    }
}

上面的查詢結果為:

{
    "took" : 1,
    "timed_out" : false,
    "_shards" : {
        "total" : 1,
        "successful" : 1,
        "skipped" : 0,
        "failed" : 0
    },
    "hits" : {
        "total" : {
            "value" : 6,
            "relation" : "eq"
        },
        "max_score" : null,
        "hits" : [ ]
    },
    "aggregations" : {
        "age_range" : {
            "buckets" : [
                {
                    "key" : "20.0-22.0",
                    "from" : 20.0,
                    "to" : 22.0,
                    "doc_count" : 1,
                    "age_min" : {
                        "value" : 20.0
                    },
                    "age_avg" : {
                        "value" : 20.0
                    },
                    "age_max" : {
                        "value" : 20.0
                    }
                },
                {
                    "key" : "22.0-25.0",
                    "from" : 22.0,
                    "to" : 25.0,
                    "doc_count" : 1,
                    "age_min" : {
                        "value" : 22.0
                    },
                    "age_avg" : {
                        "value" : 22.0
                    },
                    "age_max" : {
                        "value" : 22.0
                    }
                },
                {
                    "key" : "25.0-30.0",
                    "from" : 25.0,
                    "to" : 30.0,
                    "doc_count" : 3,
                    "age_min" : {
                        "value" : 25.0
                    },
                    "age_avg" : {
                        "value" : 26.333333333333332
                    },
                    "age_max" : {
                        "value" : 28.0
                    }
                }
            ]
        }
    }
}

Filters 聚合 (multi-bucket)

使用 Filter 聚合定義一個多存盤桶聚合,每個存盤桶都與一個過濾器相關,每個存盤桶將收集與其關聯的過濾器相匹配的所有檔案,

在上面我們使用 Range 將資料拆分成了不同的 Bucket,但是這種方式只適合欄位為數字的欄位,我們可以使用 Filter 聚合來對非數字欄位來建立不同的 Bucket,

GET twitter/_search
{
    "size": 0,
    "aggs": {
        "city_filters": {
            "filters": {
                "filters": {
                    "beijing": {
                        "match":{
                            "city":"北京"
                        }
                    },
                    "shanghai":{
                        "match":{
                            "city":"上海"
                        }
                    }
                }
            }
        }
    }
}

上面的查詢結果顯示有5個北京的檔案,一個上海的檔案,并且每個filter都有自己的名字:

{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 6,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : [ ]
  },
  "aggregations" : {
    "city_filter" : {
      "buckets" : {
        "beijing" : {
          "doc_count" : 5
        },
        "shanghai" : {
          "doc_count" : 1
        }
      }
    }
  }
}

Filter 聚合 (single-bucket)

在當前檔案背景關系中定義與指定過濾器匹配的所有檔案的單個存盤桶,通常將用于將當前聚合背景關系縮小到一組特定的檔案,

查詢城市為 北京 的檔案,并求平均年齡、最大以及最小年齡:

GET twitter/_search
{
  "size":0,
  "aggs": {
    "agg_filter": {
      "filter": {
        "match":{
          "city":"北京"
        }
      },
      "aggs": {
        "age_avg": {
          "avg": {
            "field": "age"
          }
        },
        "avg_max":{
          "max": {
            "field": "age"
          }
        },
        "avg_min":{
          "min": {
            "field": "age"
          }
        }
      }
    }
  }
}

查詢結果為:

{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 6,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : [ ]
  },
  "aggregations" : {
    "agg_filter" : {
      "doc_count" : 5,
      "avg_min" : {
        "value" : 20.0
      },
      "avg_max" : {
        "value" : 30.0
      },
      "age_avg" : {
        "value" : 24.6
      }
    }
  }
}

Date Range 聚合 (multi-bucket)

專用于日期值的范圍聚合,此聚合與正常范圍聚合之間的主要區別是,from和to值可以用Date Math運算式表示,而且還可以指定回傳from和to回應欄位的日期格式,

注意:對于每個范圍,此聚合包括from值,排除to值,

根據生日范圍查詢檔案:

GET twitter/_search
{
  "size": 0,
  "aggs": {
    "birthday_range": {
      "date_range": {
        "field": "birthday",
        "format": "yyyy-MM-dd", 
        "ranges": [
          {
            "from": "1989-04-01",
            "to": "1997-04-01"
          },
          {
            "from": "1994-04-01",
            "to": "1999-04-01"
          }
        ]
      }
    }
  }
}

查詢結果:

{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 6,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : [ ]
  },
  "aggregations" : {
    "birthday_range" : {
      "buckets" : [
        {
          "key" : "1989-04-01-1997-04-01",
          "from" : 6.07392E11,
          "from_as_string" : "1989-04-01",
          "to" : 8.598528E11,
          "to_as_string" : "1997-04-01",
          "doc_count" : 4
        },
        {
          "key" : "1994-04-01-1999-04-01",
          "from" : 7.651584E11,
          "from_as_string" : "1994-04-01",
          "to" : 9.229248E11,
          "to_as_string" : "1999-04-01",
          "doc_count" : 2
        }
      ]
    }
  }
}

Terms 聚合 (multi-bucket)

基于多桶值源的聚合,其中動態構建桶-每個唯一值一個,

可以根據 terms 聚合查詢關鍵字出現的頻率,下面我們查詢在所有檔案中出現 happy birthday 關鍵字并按照城市進行分類:

GET twitter/_search
{
  "query": {
    "match": {
      "message": "happy birthday"
    }
  },
  "size": 0, 
  "aggs": {
    "city_terms": {
      "terms": {
        "field": "city.keyword",
        "size": 10,
        "order": {
          "_count": "asc"
        }
      }
    }
  }
}

size=10 指的是排名前十的城市,并通過 doc_count 進行排序,聚合的結果為:

{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 3,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : [ ]
  },
  "aggregations" : {
    "city_terms" : {
      "doc_count_error_upper_bound" : 0,
      "sum_other_doc_count" : 0,
      "buckets" : [
        {
          "key" : "上海",
          "doc_count" : 1
        },
        {
          "key" : "北京",
          "doc_count" : 2
        }
      ]
    }
  }
}

histogram 聚合

基于多桶值源的匯總,可以應用于從檔案中提取數值或數值范圍值,根據值動態構建固定大小(也稱為間隔)的存盤桶,

GET twitter/_search
{
  "size": 0,
  "aggs": {
    "age_histogram": {
      "histogram": {
        "field": "age",
        "interval": 2
      }
    }
  }
}
  • interval : 間隔為2

回傳結果:

{
  "took" : 3,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 6,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : [ ]
  },
  "aggregations" : {
    "age_histogram" : {
      "buckets" : [
        {
          "key" : 20.0,
          "doc_count" : 1
        },
        {
          "key" : 22.0,
          "doc_count" : 1
        },
        {
          "key" : 24.0,
          "doc_count" : 1
        },
        {
          "key" : 26.0,
          "doc_count" : 1
        },
        {
          "key" : 28.0,
          "doc_count" : 1
        },
        {
          "key" : 30.0,
          "doc_count" : 1
        }
      ]
    }
  }
}

轉載請註明出處,本文鏈接:https://www.uj5u.com/shujuku/286087.html

標籤:其它

上一篇:解決動態庫的符號沖突

下一篇:kapacitor的安裝及部分常用命令

標籤雲
其他(157675) Python(38076) JavaScript(25376) Java(17977) C(15215) 區塊鏈(8255) C#(7972) AI(7469) 爪哇(7425) MySQL(7132) html(6777) 基礎類(6313) sql(6102) 熊猫(6058) PHP(5869) 数组(5741) R(5409) Linux(5327) 反应(5209) 腳本語言(PerlPython)(5129) 非技術區(4971) Android(4554) 数据框(4311) css(4259) 节点.js(4032) C語言(3288) json(3245) 列表(3129) 扑(3119) C++語言(3117) 安卓(2998) 打字稿(2995) VBA(2789) Java相關(2746) 疑難問題(2699) 细绳(2522) 單片機工控(2479) iOS(2429) ASP.NET(2402) MongoDB(2323) 麻木的(2285) 正则表达式(2254) 字典(2211) 循环(2198) 迅速(2185) 擅长(2169) 镖(2155) 功能(1967) .NET技术(1958) Web開發(1951) python-3.x(1918) HtmlCss(1915) 弹簧靴(1913) C++(1909) xml(1889) PostgreSQL(1872) .NETCore(1853) 谷歌表格(1846) Unity3D(1843) for循环(1842)

熱門瀏覽
  • GPU虛擬機創建時間深度優化

    **?桔妹導讀:**GPU虛擬機實體創建速度慢是公有云面臨的普遍問題,由于通常情況下創建虛擬機屬于低頻操作而未引起業界的重視,實際生產中還是存在對GPU實體創建時間有苛刻要求的業務場景。本文將介紹滴滴云在解決該問題時的思路、方法、并展示最終的優化成果。 從公有云服務商那里購買過虛擬主機的資深用戶,一 ......

    uj5u.com 2020-09-10 06:09:13 more
  • 可編程網卡芯片在滴滴云網路的應用實踐

    **?桔妹導讀:**隨著云規模不斷擴大以及業務層面對延遲、帶寬的要求越來越高,采用DPDK 加速網路報文處理的方式在橫向縱向擴展都出現了局限性。可編程芯片成為業界熱點。本文主要講述了可編程網卡芯片在滴滴云網路中的應用實踐,遇到的問題、帶來的收益以及開源社區貢獻。 #1. 資料中心面臨的問題 隨著滴滴 ......

    uj5u.com 2020-09-10 06:10:21 more
  • 滴滴資料通道服務演進之路

    **?桔妹導讀:**滴滴資料通道引擎承載著全公司的資料同步,為下游實時和離線場景提供了必不可少的源資料。隨著任務量的不斷增加,資料通道的整體架構也隨之發生改變。本文介紹了滴滴資料通道的發展歷程,遇到的問題以及今后的規劃。 #1. 背景 資料,對于任何一家互聯網公司來說都是非常重要的資產,公司的大資料 ......

    uj5u.com 2020-09-10 06:11:05 more
  • 滴滴AI Labs斬獲國際機器翻譯大賽中譯英方向世界第三

    **桔妹導讀:**深耕人工智能領域,致力于探索AI讓出行更美好的滴滴AI Labs再次斬獲國際大獎,這次獲獎的專案是什么呢?一起來看看詳細報道吧! 近日,由國際計算語言學協會ACL(The Association for Computational Linguistics)舉辦的世界最具影響力的機器 ......

    uj5u.com 2020-09-10 06:11:29 more
  • MPP (Massively Parallel Processing)大規模并行處理

    1、什么是mpp? MPP (Massively Parallel Processing),即大規模并行處理,在資料庫非共享集群中,每個節點都有獨立的磁盤存盤系統和記憶體系統,業務資料根據資料庫模型和應用特點劃分到各個節點上,每臺資料節點通過專用網路或者商業通用網路互相連接,彼此協同計算,作為整體提供 ......

    uj5u.com 2020-09-10 06:11:41 more
  • 滴滴資料倉庫指標體系建設實踐

    **桔妹導讀:**指標體系是什么?如何使用OSM模型和AARRR模型搭建指標體系?如何統一流程、規范化、工具化管理指標體系?本文會對建設的方法論結合滴滴資料指標體系建設實踐進行解答分析。 #1. 什么是指標體系 ##1.1 指標體系定義 指標體系是將零散單點的具有相互聯系的指標,系統化的組織起來,通 ......

    uj5u.com 2020-09-10 06:12:52 more
  • 單表千萬行資料庫 LIKE 搜索優化手記

    我們經常在資料庫中使用 LIKE 運算子來完成對資料的模糊搜索,LIKE 運算子用于在 WHERE 子句中搜索列中的指定模式。 如果需要查找客戶表中所有姓氏是“張”的資料,可以使用下面的 SQL 陳述句: SELECT * FROM Customer WHERE Name LIKE '張%' 如果需要 ......

    uj5u.com 2020-09-10 06:13:25 more
  • 滴滴Ceph分布式存盤系統優化之鎖優化

    **桔妹導讀:**Ceph是國際知名的開源分布式存盤系統,在工業界和學術界都有著重要的影響。Ceph的架構和演算法設計發表在國際系統領域頂級會議OSDI、SOSP、SC等上。Ceph社區得到Red Hat、SUSE、Intel等大公司的大力支持。Ceph是國際云計算領域應用最廣泛的開源分布式存盤系統, ......

    uj5u.com 2020-09-10 06:14:51 more
  • es~通過ElasticsearchTemplate進行聚合~嵌套聚合

    之前寫過《es~通過ElasticsearchTemplate進行聚合操作》的文章,這一次主要寫一個嵌套的聚合,例如先對sex集合,再對desc聚合,最后再對age求和,共三層嵌套。 Aggregations的部分特性類似于SQL語言中的group by,avg,sum等函式,Aggregation ......

    uj5u.com 2020-09-10 06:14:59 more
  • 爬蟲日志監控 -- Elastc Stack(ELK)部署

    傻瓜式部署,只需替換IP與用戶 導讀: 現ELK四大組件分別為:Elasticsearch(核心)、logstash(處理)、filebeat(采集)、kibana(可視化) 下載均在https://www.elastic.co/cn/downloads/下tar包,各組件版本最好一致,配合fdm會 ......

    uj5u.com 2020-09-10 06:15:05 more
最新发布
  • day02-2-商鋪查詢快取

    功能02-商鋪查詢快取 3.商鋪詳情快取查詢 3.1什么是快取? 快取就是資料交換的緩沖區(稱作Cache),是存盤資料的臨時地方,一般讀寫性能較高。 快取的作用: 降低后端負載 提高讀寫效率,降低回應時間 快取的成本: 資料一致性成本 代碼維護成本 運維成本 3.2需求說明 如下,當我們點擊商店詳 ......

    uj5u.com 2023-04-20 08:33:24 more
  • MySQL中binlog備份腳本分享

    關于MySQL的二進制日志(binlog),我們都知道二進制日志(binlog)非常重要,尤其當你需要point to point災難恢復的時侯,所以我們要對其進行備份。關于二進制日志(binlog)的備份,可以基于flush logs方式先切換binlog,然后拷貝&壓縮到到遠程服務器或本地服務器 ......

    uj5u.com 2023-04-20 08:28:06 more
  • day02-短信登錄

    功能實作02 2.功能01-短信登錄 2.1基于Session實作登錄 2.1.1思路分析 2.1.2代碼實作 2.1.2.1發送短信驗證碼 發送短信驗證碼: 發送驗證碼的介面為:http://127.0.0.1:8080/api/user/code?phone=xxxxx<手機號> 請求方式:PO ......

    uj5u.com 2023-04-20 08:27:27 more
  • 快取與資料庫雙寫一致性幾種策略分析

    本文將對幾種快取與資料庫保證資料一致性的使用方式進行分析。為保證高并發性能,以下分析場景不考慮執行的原子性及加鎖等強一致性要求的場景,僅追求最終一致性。 ......

    uj5u.com 2023-04-20 08:26:48 more
  • sql陳述句優化

    問題查找及措施 問題查找 需要找到具體的代碼,對其進行一對一優化,而非一直把關注點放在服務器和sql平臺 降低簡化每個事務中處理的問題,盡量不要讓一個事務拖太長的時間 例如檔案上傳時,應將檔案上傳這一步放在事務外面 微軟建議 4.啟動sql定時執行計劃 怎么啟動sqlserver代理服務-百度經驗 ......

    uj5u.com 2023-04-20 08:26:35 more
  • 云時代,MySQL到ClickHouse資料同步產品對比推薦

    ClickHouse 在執行分析查詢時的速度優勢很好的彌補了MySQL的不足,但是對于很多開發者和DBA來說,如何將MySQL穩定、高效、簡單的同步到 ClickHouse 卻很困難。本文對比了 NineData、MaterializeMySQL(ClickHouse自帶)、Bifrost 三款產品... ......

    uj5u.com 2023-04-20 08:26:29 more
  • sql陳述句優化

    問題查找及措施 問題查找 需要找到具體的代碼,對其進行一對一優化,而非一直把關注點放在服務器和sql平臺 降低簡化每個事務中處理的問題,盡量不要讓一個事務拖太長的時間 例如檔案上傳時,應將檔案上傳這一步放在事務外面 微軟建議 4.啟動sql定時執行計劃 怎么啟動sqlserver代理服務-百度經驗 ......

    uj5u.com 2023-04-20 08:25:13 more
  • Redis 報”OutOfDirectMemoryError“(堆外記憶體溢位)

    Redis 報錯“OutOfDirectMemoryError(堆外記憶體溢位) ”問題如下: 一、報錯資訊: 使用 Redis 的業務介面 ,產生 OutOfDirectMemoryError(堆外記憶體溢位),如圖: 格式化后的報錯資訊: { "timestamp": "2023-04-17 22: ......

    uj5u.com 2023-04-20 08:24:54 more
  • day02-2-商鋪查詢快取

    功能02-商鋪查詢快取 3.商鋪詳情快取查詢 3.1什么是快取? 快取就是資料交換的緩沖區(稱作Cache),是存盤資料的臨時地方,一般讀寫性能較高。 快取的作用: 降低后端負載 提高讀寫效率,降低回應時間 快取的成本: 資料一致性成本 代碼維護成本 運維成本 3.2需求說明 如下,當我們點擊商店詳 ......

    uj5u.com 2023-04-20 08:24:03 more
  • day02-短信登錄

    功能實作02 2.功能01-短信登錄 2.1基于Session實作登錄 2.1.1思路分析 2.1.2代碼實作 2.1.2.1發送短信驗證碼 發送短信驗證碼: 發送驗證碼的介面為:http://127.0.0.1:8080/api/user/code?phone=xxxxx<手機號> 請求方式:PO ......

    uj5u.com 2023-04-20 08:23:11 more