文章詳情頁

MySQL去重該使用distinct還是group by？

瀏覽：4日期：2023-10-15 09:04:48

前言

關于group by 與distinct 性能對比:網上結論如下，不走索引少量數據distinct性能更好，大數據量group by 性能好，走索引group by性能好。走索引時分組種類少distinct快。關于網上的結論做一次驗證。

準備階段屏蔽查詢緩存

查看MySQL中是否設置了查詢緩存。為了不影響測試結果，需要關閉查詢緩存。

show variables like ’%query_cache%’;

MySQL去重該使用distinct還是group by？

查看是否開啟查詢緩存決定于query_cache_type和query_cache_size。

方法一：關閉查詢緩存需要找到my.ini，修改query_cache_type需要修改C:ProgramDataMySQLMySQL Server 5.7my.ini配置文件，修改query_cache_type=0或2。方法二：設置query_cache_size為0，執行以下語句。

set global query_cache_size = 0;

方法三：如果你不想關閉查詢緩存，也可以在使用RESET QUERY CACHE。

現在測試環境中query_cache_type=2代表按需進行查詢緩存，默認的查詢方式是不會進行緩存，如需緩存則需要在查詢語句中加上sql_cache。

數據準備

t0表存放10W少量種類少的數據

drop table if exists t0;create table t0(id bigint primary key auto_increment,a varchar(255) not null) engine=InnoDB default charset=utf8mb4 collate=utf8mb4_bin;12345drop procedure insert_t0_simple_category_data_sp;delimiter //create procedure insert_t0_simple_category_data_sp(IN num int)beginset @i = 0;while @i < num doinsert into t0(a) value(truncate(@i/1000, 0)); set @i = @i + 1;end while;end//call insert_t0_simple_category_data_sp(100000);

t1表存放1W少量種類多的數據

drop table if exists t1;create table t1 like t0;12drop procedure insert_t1_complex_category_data_sp;delimiter //create procedure insert_t1_complex_category_data_sp(IN num int)beginset @i = 0;while @i < num doinsert into t1(a) value(truncate(@i/10, 0)); set @i = @i + 1;end while;end//call insert_t1_complex_category_data_sp(10000);

t2表存放500W大量種類多的數據

drop table if exists t2;create table t2 like t1;12drop procedure insert_t2_complex_category_data_sp;delimiter //create procedure insert_t2_complex_category_data_sp(IN num int)beginset @i = 0;while @i < num doinsert into t1(a) value(truncate(@i/10, 0)); set @i = @i + 1;end while;end//call insert_t2_complex_category_data_sp(5000000);

測試階段

驗證少量種類少數據

未加索引

set profiling = 1;select distinct a from t0;show profiles;select a from t0 group by a;show profiles;alter table t0 add index `a_t0_index`(a);

MySQL去重該使用distinct還是group by？

由此可見：少量種類少數據下，未加索引，distinct和group by性能相差無幾。

加索引

alter table t0 add index `a_t0_index`(a);

執行上述類似查詢后

MySQL去重該使用distinct還是group by？

由此可見：少量種類少數據下，加索引，distinct和group by性能相差無幾。

驗證少量種類多數據未加索引

執行上述類似未加索引查詢后

MySQL去重該使用distinct還是group by？

由此可見：少量種類多數據下，未加索引，distinct比group by性能略高，差距并不大。

加索引

alter table t1 add index `a_t1_index`(a);

執行類似未加索引查詢后

MySQL去重該使用distinct還是group by？

由此可見：少量種類多數據下，加索引，distinct和group by性能相差無幾。

驗證大量種類多數據

未加索引

SELECT count(1) FROM t2;

MySQL去重該使用distinct還是group by？

執行上述類似未加索引查詢后

MySQL去重該使用distinct還是group by？

由此可見：大量種類多數據下，未加索引，distinct比group by性能高。

加索引

alter table t2 add index `a_t2_index`(a);

執行上述類似加索引查詢后

MySQL去重該使用distinct還是group by？

由此可見：大量種類多數據下，加索引，distinct和group by性能相差無幾。

總結性能比少量種類少少量種類多大量種類多未加索引相差無幾distinct略優distinct更優加索引相差無幾相差無幾相差無幾

去重場景下，未加索引時，更偏向于使用distinct，而加索引時，distinct和group by兩者都可以使用。

總結

到此這篇關于MySQL去重該使用distinct還是group by？的文章就介紹到這了,更多相關mysql 去重distinct group by內容請搜索好吧啦網以前的文章或繼續瀏覽下面的相關文章希望大家以后多多支持好吧啦網！

MySQL 數據庫

上一條：MySQL explain獲取查詢指令信息原理及實例下一條：MySQL 事務概念與用法深入詳解

相關文章：

1. db2v8的pdf文檔資料2. MySQL之高可用集群部署及故障切換實現3. 不要忽視Oracle 10g STATSPACK新功能4. DB2 數據庫應用中使用受信任上下文（1）5. Oracle一行拆分為多行方法實例6. IDEA找不到Database的完美解決方法7. 如何遠程調用ACCESS數據庫8. 在DB2中如何實現Oracle的相關功能(一)9. MySQL基本調度策略淺析10. oracle基本概念和術語

国产综合久久一区二区三区