ReUtil
ReUtil是Hutool中用于处理正则表达式的核心工具类,提供了丰富的正则表达式操作功能,包括正则匹配、查找、替换、分割等。
# 主要功能
- 正则匹配:判断字符串是否匹配指定的正则表达式
- 正则查找:查找字符串中匹配正则表达式的部分
- 正则替换:替换字符串中匹配正则表达式的部分
- 正则分割:使用正则表达式分割字符串
- 正则提取:提取字符串中匹配正则表达式的分组
- 正则捕获:捕获字符串中匹配正则表达式的所有部分
- 正则替换:支持函数式替换
# 正则匹配
# isMatch(String regex, CharSequence content)
判断字符串是否匹配指定的正则表达式。
// 验证邮箱格式
boolean isEmail = ReUtil.isMatch("^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$", "test@example.com");
// 输出: true
// 验证手机号格式
boolean isPhone = ReUtil.isMatch("^1[3-9]\\d{9}$", "13812345678");
// 输出: true
# contains(String regex, CharSequence content)
判断字符串中是否包含匹配正则表达式的部分。
String text = "abc123def456ghi789";
boolean containsNum = ReUtil.contains("\\d+", text);
// 输出: true
boolean containsLetter = ReUtil.contains("[a-zA-Z]+", text);
// 输出: true
# 正则查找
# get(String regex, CharSequence content, int groupIndex)
获取字符串中匹配正则表达式的指定分组。
String text = "abc123def456ghi789";
String num = ReUtil.get("\\d+", text, 0);
// 输出: 123
String url = "https://www.baidu.com/s?wd=hutool";
String domain = ReUtil.get("https?://([\\w.-]+)/", url, 1);
// 输出: www.baidu.com
# getGroup1(String regex, CharSequence content)
获取字符串中匹配正则表达式的第1个分组。
String text = "name:张三,age:25";
String name = ReUtil.getGroup1("name:(\\w+),age:(\\d+)", text);
// 输出: 张三
# getGroup2(String regex, CharSequence content)
获取字符串中匹配正则表达式的第2个分组。
String text = "name:张三,age:25";
String age = ReUtil.getGroup2("name:(\\w+),age:(\\d+)", text);
// 输出: 25
# findAll(String regex, CharSequence content)
查找字符串中所有匹配正则表达式的部分,返回Matcher对象列表。
String text = "abc123def456ghi789";
List<Matcher> matchers = ReUtil.findAll("\\d+", text);
for (Matcher matcher : matchers) {
Console.log(matcher.group());
}
// 输出:
// 123
// 456
// 789
# findAllGroup0(String regex, CharSequence content)
查找字符串中所有匹配正则表达式的部分,返回匹配的字符串列表。
String text = "abc123def456ghi789";
List<String> nums = ReUtil.findAllGroup0("\\d+", text);
// 输出: ["123", "456", "789"]
# findAllGroup1(String regex, CharSequence content)
查找字符串中所有匹配正则表达式的第1个分组,返回分组的字符串列表。
String text = "name:张三,age:25;name:李四,age:30";
List<String> names = ReUtil.findAllGroup1("name:(\\w+),age:(\\d+)", text);
// 输出: ["张三", "李四"]
# 正则替换
# replaceAll(String regex, CharSequence content, String replacement)
替换字符串中所有匹配正则表达式的部分。
String text = "abc123def456ghi789";
String result = ReUtil.replaceAll("\\d+", text, "*");
// 输出: abc*def*ghi*
# replaceFirst(String regex, CharSequence content, String replacement)
替换字符串中第一个匹配正则表达式的部分。
String text = "abc123def456ghi789";
String result = ReUtil.replaceFirst("\\d+", text, "*");
// 输出: abc*def456ghi789
# replaceAll(String regex, CharSequence content, Function<Matcher, String> replacer)
使用函数式替换字符串中所有匹配正则表达式的部分。
String text = "abc123def456ghi789";
String result = ReUtil.replaceAll("\\d+", text, matcher -> {
int num = Integer.parseInt(matcher.group());
return String.valueOf(num * num);
});
// 输出: abc15129def207936ghi622521
# replaceFirst(String regex, CharSequence content, Function<Matcher, String> replacer)
使用函数式替换字符串中第一个匹配正则表达式的部分。
String text = "abc123def456ghi789";
String result = ReUtil.replaceFirst("\\d+", text, matcher -> {
int num = Integer.parseInt(matcher.group());
return String.valueOf(num * num);
});
// 输出: abc15129def456ghi789
# 正则分割
# split(String regex, CharSequence content)
使用正则表达式分割字符串。
String text = "Hello World Hutool";
String[] parts = ReUtil.split("\\s+", text);
// 输出: ["Hello", "World", "Hutool"]
# split(String regex, CharSequence content, int limit)
使用正则表达式分割字符串,指定分割次数。
String text = "a,b,c,d,e";
String[] parts = ReUtil.split(",", text, 3);
// 输出: ["a", "b", "c,d,e"]
# 正则捕获
# extractMulti(String regex, CharSequence content)
提取字符串中匹配正则表达式的所有分组,返回二维数组。
String text = "name:张三,age:25;name:李四,age:30";
String[][] result = ReUtil.extractMulti("name:(\\w+),age:(\\d+)", text);
// 输出: [["name:张三,age:25", "张三", "25"], ["name:李四,age:30", "李四", "30"]]
# extractMulti(String regex, CharSequence content, int groupCount)
提取字符串中匹配正则表达式的指定数量的分组,返回二维数组。
String text = "name:张三,age:25;name:李四,age:30";
String[][] result = ReUtil.extractMulti("name:(\\w+),age:(\\d+)", text, 2);
// 输出: [["张三", "25"], ["李四", "30"]]
# 正则替换为文件
# extractMultiToFile(String regex, File file, int groupCount, File destFile)
从文件中提取匹配正则表达式的分组,并写入到目标文件中。
File sourceFile = new File("source.txt");
File destFile = new File("dest.txt");
ReUtil.extractMultiToFile("name:(\\w+),age:(\\d+)", sourceFile, 2, destFile);
# 常用正则表达式
ReUtil可以配合RegexPool使用,RegexPool内置了大量常用正则表达式,如邮箱、手机号、URL等。
// 使用内置的邮箱正则
boolean isEmail = ReUtil.isMatch(RegexPool.EMAIL, "test@example.com");
// 输出: true
// 使用内置的手机号正则
boolean isPhone = ReUtil.isMatch(RegexPool.MOBILE, "13812345678");
// 输出: true
// 使用内置的URL正则
boolean isUrl = ReUtil.isMatch(RegexPool.URL, "https://www.baidu.com");
// 输出: true
// 使用内置的身份证正则
boolean isIdCard = ReUtil.isMatch(RegexPool.ID_CARD, "110101199001011234");
// 输出: true
# 正则表达式缓存
ReUtil内置了正则表达式缓存机制,会缓存编译后的正则表达式,提高性能。缓存会在首次调用时自动创建,后续调用会直接从缓存中获取,避免重复的正则表达式编译。
// 第一次调用,会编译正则表达式并缓存
boolean isEmail1 = ReUtil.isMatch(RegexPool.EMAIL, "test1@example.com");
// 第二次调用,会直接从缓存中获取编译后的正则表达式
boolean isEmail2 = ReUtil.isMatch(RegexPool.EMAIL, "test2@example.com");
# 注意事项
- 正则表达式语法:使用Java正则表达式语法,注意转义字符的使用
- 分组索引:分组索引从0开始,0表示整个匹配的字符串
- 性能考虑:对于频繁使用的正则表达式,建议使用缓存机制
- 贪婪匹配:Java正则表达式默认是贪婪匹配,可以使用?来切换为非贪婪匹配
- 字符集:注意正则表达式中的字符集,如\d表示数字,\w表示字母、数字、下划线
- 边界匹配:使用^和$来匹配字符串的开始和结束
- 量词:使用*、+、?、{n}、{n,}、{n,m}等来表示量词
# 总结
ReUtil提供了丰富的正则表达式操作功能,涵盖了正则匹配、查找、替换、分割、提取等多个方面。通过使用ReUtil,可以大大简化Java中的正则表达式操作,提高开发效率。
无论是文本验证、文本提取、文本替换还是文本分割,ReUtil都能提供简洁高效的解决方案。其内置的常用正则表达式和缓存机制也提高了开发效率和性能。