Python 文件操作实战:安全整理下载目录
用 pathlib 写安全的文件整理脚本:预览模式、防覆盖、递归扫描与原子替换。
文件操作的脚本大多死在同一个问题上:没想清楚「如果文件已经存在」「如果目录不存在」「如果权限不够」这三种情况。这篇文章用 pathlib 写一个「整理下载目录」的脚本,重点是预览模式、防覆盖和错误处理——这些才是生产级文件脚本和玩具脚本的分水岭。
#pathlib 比 os.path 好在哪
pathlib 把路径变成对象,. 和 .. 的解析、分隔符、拼接都自动处理,跨平台一致:
from pathlib import Path
home = Path.home()
downloads = home / "Downloads"
target = downloads / "images" / "a.png"
print(downloads.exists()) # 布尔判断
print(target.suffix) # .png
print(target.stem) # a
print(target.with_suffix(".jpg")) # images/a.jpg注意 / 运算符是拼接,不要写成字符串相加。Path 对象之间、Path 和字符串混用时,先统一成 Path 再操作。
#分类整理:先预览,后执行
直接 shutil.move 的脚本,一旦正则写错,文件就进了错误目录。我的做法是默认「只打印要做什么」,加 --apply 才真动手:
from __future__ import annotations
import argparse
import shutil
from pathlib import Path
CATEGORIES = {
".jpg": "images", ".jpeg": "images", ".png": "images", ".gif": "images",
".pdf": "documents", ".docx": "documents", ".txt": "documents",
".zip": "archives", ".tar.gz": "archives", ".rar": "archives",
}
def unique_destination(path: Path) -> Path:
if not path.exists():
return path
for index in range(1, 10000):
candidate = path.with_name(f"{path.stem}_{index}{path.suffix}")
if not candidate.exists():
return candidate
raise RuntimeError(f"无法为 {path.name} 生成唯一名称")
def organize(folder: Path, apply: bool) -> None:
for source in folder.iterdir():
if not source.is_file() or source.name.startswith("."):
continue
category = CATEGORIES.get(source.suffix.lower(), "others")
destination = unique_destination(folder / category / source.name)
print(f"{source.name} -> {destination.relative_to(folder)}")
if apply:
destination.parent.mkdir(parents=True, exist_ok=True)
shutil.move(str(source), str(destination))几个细节:
CATEGORIES.get(source.suffix.lower(), "others"):.JPG和.jpg都能命中,认不出的进others,别让未知类型抛异常。unique_destination:同名文件自动加_1、_2,而不是覆盖。覆盖是文件脚本最危险的默认行为。- 跳过后缀名是空字符串的隐藏文件(
.DS_Store、~临时文件),防止把系统文件挪走。 - 打印的是相对路径,方便人核对;日志要可读,别打一长串绝对路径。
#递归扫描:glob 和 rglob 的区别
iterdir() 只看当前目录一层。要递归,用 glob(带通配符)或 rglob(递归匹配):
for pdf in folder.rglob("*.pdf"):
print(pdf)rglob 会走进所有子目录,包括 __pycache__ 这种,也可能跟着符号链接走。扫描前先想清楚要不要 follow symlink,Path.is_symlink() 检查一下:
for item in folder.rglob("*"):
if item.is_symlink():
continue # 跳过符号链接,防止循环引用
if item.is_file():
...#错误处理:谁出错就记谁,别让整个脚本崩掉
一个文件权限不够,不应该中断整个整理任务:
def organize(folder: Path, apply: bool) -> None:
errors = []
for source in folder.iterdir():
try:
...
except (PermissionError, OSError) as e:
errors.append((source.name, str(e)))
if errors:
print("以下文件处理失败:")
for name, msg in errors:
print(f" {name}: {msg}")
raise SystemExit(1)权限错误的典型场景:目录属于别的用户、文件只读、磁盘满(OSError: [Errno 28] No space left on device)。磁盘满时 shutil.move 可能把源文件写坏一半——跨文件系统移动时尤其危险,移动过程中断电或满盘,源文件可能残缺。
#跨文件系统移动和原子替换
shutil.move 在同一文件系统内是改名(快),跨文件系统则是先复制再删除(慢、且非原子)。如果你要「替换已存在的文件」,别用 shutil.move 直接盖,用 os.replace——它是原子操作,要么新内容完全就位,要么原样不动:
import os
os.replace(temp_path, final_path) # 原子替换,进程崩溃也不会留半个文件先写临时文件、再 os.replace 到目标名,这是「写配置文件」「更新导出文件」的标准安全写法,我在脚本里一直用。
#经验
文件脚本最大的风险不是逻辑错,而是人在没有预览的情况下跑了 --apply。所以我的习惯:
- 默认预览,
--apply才动手; - 目标路径永远生成唯一名,绝不覆盖;
- 每个文件单独 try/except,记录失败不中断;
- 动重要目录前先
du -sh看一眼总量,心里有数再跑。
另外 Path.read_text() / write_text() 记得显式传 encoding="utf-8",否则在 Windows 上会按 GBK 读写,中文直接乱码——这个坑我帮同事排查过整整一个下午。
写于 2025 年 10 月 12 日
- 栏目
- 技术文章
- 约
- 3.7 分钟
- 字数
- 3K
- 阅读
- 206
本文为原创记录,转载请注明出处。如果这篇替你省了时间,欢迎留言说说你踩到的坑。
同题 · related
留言 · remarks
00 条还没有留言,来说点什么吧。